Skip to content
José Luis Delgado
AI Systems Evaluation Automation

Agentic Evaluation Harness

An evaluation scaffold for automated research workflows: scenario design, trace capture, failure review, and reproducible comparisons between agents.

2026
Evaluation engineering

Problem

Agentic tooling can look impressive while hiding the hard questions: observed state, changed files, failures, and reproducibility status.

Technical record

The harness treats automated work as an experiment. It records scenarios, tool calls, outputs, failures, and review records so comparisons are inspectable rather than anecdotal.

Evidence

Evidence trails expose quality, regressions, and failure modes across runs.