Technology · Agentic → Evaluations · Open Source · wiki:deep
Agent Evaluation (AWS Labs) is a generative AI-powered framework for testing virtual agents. Internally it runs an LLM evaluator agent that orchestrates multi-turn conversations with your target agent and scores responses during the conversation — not a static prompt-string checker.
It ships built-in targets for Amazon Bedrock, Amazon Q Business, and SageMaker, plus custom targets; supports concurrent multi-turn runs, hooks for integration tests, and CI/CD wiring so agent changes can fail the pipeline when conversations regress.
Committed criteria need scored agent conversations, not vibe demos. This pack is the AWS-shaped harness for “dataset/scenario → talk to agent → metric → gate merge.” Use it when the install already lives on Bedrock/Q/SageMaker or when you need a conversation-orchestrating evaluator rather than only RAG metric libraries (Ragas) or declarative YAML evals (Promptfoo).
Wire scores to cases so a failing eval is attributable; never treat the harness as process authority.
Related: abm-agent-behavior-mining · deepeval · hal-the-holistic-agent-leaderboard · topics/32-eval-harnesses
Scroll inside the canvas to pan
science/sources mix in HAL/SWE-bench contrast files — Evidence rows cite what exists; deepen sources when claiming Bedrock-specific facts.Catalog backlinks — what points here (wiki “what links here”).
Claims below are backed by science sources on disk.
Conversational / harness eval framing
HAL harness related · MODERATE
“Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities.”
Benchmark contrast (coding agents)
SWE-bench: Can Language Models Resolve Real-World GitHub Issues? · MODERATE
“Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue.”
Peer contrast vs RAG metrics
Component eval contrast · MODERATE
“We find real-world software engineering to be a rich, sustainable, and challenging testbed for evaluating the next generation of language models.”
Features and peers linked from the catalog map — not a second product surface.
No DM vendor crosswalk edges yet.
Primary repo github.com/awslabs/agent-evaluation · Open Source
technologies/agent-evaluation/README.md
9 tags · 14 out · 15 in · 3 artifacts · 1 gaps · 29 corpus docs
map_edge · 7
tech_features · 1
tech_quote · 7
tech_readme · 1
tech_science_source · 3
tech_section · 10