The exam: the USA Hockey Boys National 17 Festival, July 7 to 13 in Amherst, New York. Twelve district teams. 216 of the best American players born in 2009, selected by the federation's own evaluators, playing five games each on the road to the Hlinka Gretzky selection camp. The question we tested: does a model that has never watched these kids skate agree with the people whose job is to pick them?
Before the results, the ground rules, because a self-grade is worthless without them.
We graded against the full field. All 216 selections, every skater stat line including the zeros, every goalie. No cherry-picking the kids who scored. Our grades were snapshotted before the festival began, so nothing here is fitted after the fact. And five games is a small, noisy sample: linemates matter, defensemen barely score, district quotas guarantee 18 players per district regardless of depth, and a development camp is not optimized for winning. Keep all of that in mind. We did.
All but a handful of the 216 selections were already in our database, ranked, before USA Hockey named them. The selected skaters sit in the top 2 to 3 percent of our entire US 2009 board: the median selected forward ranks #99 out of 5,169 American forwards we score in that birth year, the median defenseman #81 out of 2,494. Exact top-N overlap: 55 percent for forwards, 50 percent for defense, 25 percent for goalies. On skaters, a model reading nothing but data and the federation's scouts watching thousands of games converge on largely the same names.
Result 2: our grades predicted production
We bucketed all 116 matched forwards by our pre-camp 0 to 99 grade and checked what they did on the ice in Amherst:
- Grade 85+: 23 forwards, 3.96 points on average, 61 percent produced at 0.8 points per game or better.
- Grade 75 to 84: 26 forwards, 3.04 average, 38 percent.
- Grade below 75: 67 forwards, 2.51 average, 28 percent.
A clean staircase. Our top-graded forwards were more than twice as likely to produce as our lowest tier. The rank correlation is about 0.30, significant at p ≈ 0.001, and honestly modest, which is exactly what a season-long grade tested against a five-game tournament should look like. Anyone selling you a stronger number off five games is selling you noise.
Result 3: the names, both kinds
The festival scoring leader was Sawyer Schmidt of New York: 11 points in 5 games. Our model graded him 88 before camp. Our highest-graded skater in attendance, Brooks DeMars at 91, finished top three in scoring with 8. Nolan Snyder, graded 89, put up 7. Mark Pape, 89, scored 6.
Now the other column, because this is a report card and not a highlight reel. Luke Warrener carried our 90 and went pointless. Landon Jackman, 89, also zero. Michael Tang, 88, one point. Some of that is five-game variance, and one case shows why: Declan Wotton, graded 90, managed just 2 points, and both were game-winning goals. Small samples hide as much as they show. But we log the misses next to the hits, every time, because that is the whole point of grading in public.
Result 4: our goalie model needs work, and we are saying so first
Our top-graded goalie at the festival, Hagan Bach at 88, finished second statistically: 0.80 goals against, .949 save percentage. A real hit. But the best goalie in Amherst was Julian Herrera, .955 and a 0.77, and our model had him at 66. Another netminder our model graded near the floor posted a .923. Goalie agreement with USA Hockey's selectors was 25 percent against 50 to 55 for skaters. Goaltending is the hardest position to evaluate from data, every analytics group in hockey knows it, and our numbers now quantify exactly how far behind our goalie instrument is. It moves to the top of the fix list.

One structural note on the disagreements. USA Hockey allocates every district 18 players regardless of depth. A deep district leaves top-100 kids at home while a thin one sends players ranked far lower, and that quota design explains a real share of the roughly 45 percent of selections that sat outside our top-N. We published a framework for measuring exactly this earlier this year, the
Selection Bias Index, and the festival fits the same pattern: much of the scout-versus-model gap is structure, not talent evaluation.
What happens now
The festival flows back into the rankings. Selection to a national camp earns points on our board, and performance in Amherst earns more, the same way we score international tournaments. The board will move, some kids will jump hundreds of ProdigyPoints, and we will publish the biggest movers. Then the next exam grades us again: the Hlinka camp, the season, and eventually the drafts these 2009s enter.
That is the loop. Model, prediction, public grade, correction, repeat. 190,000+ players across 91 countries, and every one of them deserves an evaluation system willing to show its report card.
See you at the next exam.
POWER TO THE PLAYERS!
Philippe ⛓️