Technology · Agentic → Evaluations · Open Source · wiki:deep
HAL is Princeton’s Holistic Agent Leaderboard ecosystem: a public leaderboard plus a standardized hal-eval harness for reproducible agent runs across benchmarks (SWE-bench Verified, USACO, AppWorld, CORE-bench, τ-bench, …). The harness supports local conda/Docker and Azure VMs, Weave logging/cost tracking, and encrypted trace upload to Hugging Face for leaderboard submission.
Important (primary README, 2026): leaderboard results are no longer updated through this harness; the repo is archived for historical reference while work shifts to agent reliability. Public results remain on the HAL site and Reliability Dashboard.
Before locking a coding/tool agent stack, you need external baselines, not only private golden sets. HAL is the research pointer for multi-benchmark agent comparison and published traces — use historical harness docs to understand how public scores were produced. It is not a production monitor and not a substitute for install-specific CI goldens (Promptfoo/DeepEval).
hal from the harness repo (historically recursive clone + conda). hal-eval runs agents against a chosen benchmark with Weave logging. hal-upload sends encrypted traces for leaderboard integration. Related: abm-agent-behavior-mining · agent-evaluation · deepeval · topics/32-eval-harnesses
Scroll inside the canvas to pan
Catalog backlinks — what points here (wiki “what links here”).
Claims below are backed by science sources on disk.
Founding paper / HAL harness
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation · STRONG
“We introduce the Holistic Agent Leaderboard (HAL) to address these challenges.”
HAL harness / leaderboard role
HAL repo · MODERATE
“SWE-bench, an evaluation framework consisting of 2,294 software engineering problems drawn from real GitHub issues and corresponding pull requests across 12 popular Python repositories.”
Coding-benchmark contrast
SWE-bench · MODERATE
“Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities.”
Eval-peer contrast
RAGAS contrast · MODERATE
“Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue.”
Features and peers linked from the catalog map — not a second product surface.
No DM vendor crosswalk edges yet.
Primary repo github.com/princeton-pli/hal-harness · Open Source
technologies/hal-the-holistic-agent-leaderboard/README.md
9 tags · 14 out · 15 in · 3 artifacts · 1 gaps · 30 corpus docs
map_edge · 7
tech_features · 1
tech_quote · 8
tech_readme · 1
tech_science_source · 3
tech_section · 10