The part I’d be most careful with is that a self-improving eval loop can make a bad product intuition look increasingly rigorous. If Claude suggests evals from your existing traces, then your eval suite is downstream of whatever your current users, workflows, and product assumptions already expose. That is useful, but also dangerous: you can end up optimizing the agent around the visible failure modes while missing the higher-order question of whether the agent is solving the right job in the first place.
Traces only contain the jobs your agent already attempted, so an eval suite built from them is a closed loop: it can prove you're getting better answers to a question nobody checked was the right one. That's the failure I wrote about in https://thesynthesisai.substack.com/p/the-right-answer-to-the-wrong-question, where rigor on the numbers hides the cheaper, costlier error upstream. The fix isn't more evals from traces, it's a separate eval that asks whether the job itself should exist.
The hire-on-the-spot eval moment Aparna names matches a shift I track in PM interview loops at Indian GCCs. A year ago, near-zero GenAI PM JDs in my dataset included a hands-on eval task in the loop. Today it is common, and the shortlisted PMs are no longer the ones with the cleanest roadmap doc but the ones who treat eval-running as a craft. What conversion rate are you seeing from bundle students who actually build evals vs the ones who only read about them?
Zia. AI career strategist for Indian professionals. itszia.ai
What a generous walk though, beautifully articulated. This is such underrated content.
Appreciate it Joe!
The part I’d be most careful with is that a self-improving eval loop can make a bad product intuition look increasingly rigorous. If Claude suggests evals from your existing traces, then your eval suite is downstream of whatever your current users, workflows, and product assumptions already expose. That is useful, but also dangerous: you can end up optimizing the agent around the visible failure modes while missing the higher-order question of whether the agent is solving the right job in the first place.
Good point
Traces only contain the jobs your agent already attempted, so an eval suite built from them is a closed loop: it can prove you're getting better answers to a question nobody checked was the right one. That's the failure I wrote about in https://thesynthesisai.substack.com/p/the-right-answer-to-the-wrong-question, where rigor on the numbers hides the cheaper, costlier error upstream. The fix isn't more evals from traces, it's a separate eval that asks whether the job itself should exist.
The hire-on-the-spot eval moment Aparna names matches a shift I track in PM interview loops at Indian GCCs. A year ago, near-zero GenAI PM JDs in my dataset included a hands-on eval task in the loop. Today it is common, and the shortlisted PMs are no longer the ones with the cleanest roadmap doc but the ones who treat eval-running as a craft. What conversion rate are you seeing from bundle students who actually build evals vs the ones who only read about them?
Zia. AI career strategist for Indian professionals. itszia.ai