In our first blind-graded benchmark, a correct route beat an unassisted agent 10/10 — but a wrong-but-related route lost 0/2, to agents with no help at all. Why confidently-wrong retrieval is the expensive failure class, and how semantic matching plus a refusal threshold flipped both failures into wins.