You're right, and thanks for actually running it. I checked on my side too: 997/1000 at seed 10001, 258 template + impact + tail combos in total, none of them with two labels. So 72.7% / 97.1% is mostly the heads remembering combos, not generalizing.
Ran both tests.
One template per category held out (5 of 20), trained on the rest, 1,000 tickets from the held-out templates only:
- category: sure on 100%, right on 66.1%
- urgency: sure on 44.8%, right on 96.2% of those
- needs_human: sure on 96%, right on 87.1% of those
- all three sure on 42.6%, all three right on 51.9% of those
Channel and plan flipped on the published heads: category changes on 2.7%, needs_human on 3.5%, urgency on 7.1%.
So 72.7% doesn't survive the first one. And category stays fully sure while it's wrong a third of the time, because the threshold is picked on a random holdout of the same templates and says nothing about a template it never saw. urgency's threshold held up, category's didn't.
On real traffic that's what check sampling is for: a share of local answers still goes to the teacher and the site drops back to shadow when agreement falls. But the demo eval should have shown this. I'll put both results on the model card, call the eval unseen states instead of unseen tickets, and fix the money back cue in the rule. Same 21 of 368 payouts here btw, good catch.