Published material
I read Viv’s “Towards Automating Eval Engineering”, about a skill that builds evals from an agent repository and its traces. What stands out is that it does not try to one-shot the eval. It interviews the user, builds a Harbor task, and then checks both the agent and verifier trajectories for shortcuts. I like the idea of turning real failures into durable tests; I just want the loop to stay small enough that the eval system does not become another thing to constantly maintain.
Loading published artifact…
↧towards-automating-eval-engineering-reading-note.html12 KB · published source