Evaluation structures shape the work
A metric is not only a report; it changes which behavior gets selected and repeated.
Current explanation
When a workflow rewards visible completion, agents learn to optimize visible completion. Stronger evaluation names the behavior wanted, preserves negative results, checks the target path, and makes uncertainty legible. The structure should select for judgment rather than merely more artifacts or longer sessions.
Lesson path
- 01
Name the objective
currentWrite the behavior the evaluation should select.
- 02
Predict gaming
nextList the cheapest behavior that would score without helping.
- 03
Add a countermeasure
optionalRequire one falsifiable receipt or useful non-result.
Open questions
- Which parts of judgment resist a single scalar score?
Selected Q&A
Why not maximize coverage?
Coverage can reward shallow breadth unless each specimen must survive evidence, ownership, and relevance checks.
Next actions
- Audit one dashboard for the behavior its labels encourage.
Sources
- (E)valuation Structures
Full Scrapbook essay and its revision trail.
Revision trail
· learning:evaluation-structures-shape-the-work@r1
Compressed the essay into a revisitable learning record.