A twenty-two-minute win and a fifty-eight-minute win give my original training objective the same label: one.

Duration was already in the data. Perhaps using it would help the model learn something the binary result discarded. That was worth investigating, though a short game is not an annotation saying "the draft was good." Execution, abandons, hero choices, and the pace of the match can all affect when it ends.

The preceding attempt to estimate a universal performance ceiling hadn't settled how much room remained. A cross-fitted lookup doing worse than a baseline didn't mean the opportunity was negative; it meant that estimator was noisy. And the hybrid's +0.00086 concordance margin from Part 11 could not tell me how many extra games a player would win. It was an ordering statistic, not a policy effect.

So this was a new question about supervision, not a way to calculate the value of perfect drafting.

Check the Target Before Spending the GPU

I defined a family of targets blending the binary outcome with a duration-weighted version. Faster games received more weight; a dial called alpha controlled how much of that weighting entered the target. Alpha zero recovered the original label.

Before training, I counted pair effects on two halves of the data and compared them. The gate used a count-weighted correlation between the halves' estimated counter effects, after subtracting additive hero terms. Repeating the partition measured how that statistic varied with the split.

This was deliberately narrower than asking whether duration predicted the winner. Radiant's win rate differed strongly between short and long games, so a new target could become easier to predict simply by sharpening a side effect. Removing additive terms addressed that concern within the chosen estimator; it did not eliminate every possible confound.

On 22.2 million training matches, the mean split-half correlation rose from 0.9261 for the binary outcome to 0.9374 at alpha 0.75.

The change cleared the gate's recorded repeatability criterion. Synthetic checks had also tested both a no-pair-effect case and a case where informative duration should help.

That made the target a candidate for training. It did not establish that it was a better label for predicting wins, or that the correlation was a fraction of useful information the model would recover. More repeatable can mean more repeatable information about a different property of the game.

Training Answered a Different Question

I trained the duration-weighted target against the binary control at three paired seeds, keeping evaluation tied to win/loss rather than changing the exam to reward the new label.

The chosen metrics deteriorated. A representative counter-ranking partial correlation, measured against training pair rates, fell from about 0.95 to 0.73, and the direction was consistent across seeds. On validation matchups, the mean change in the historical top-one outcome score was about -0.59 percentage points. That score concerns observed outcomes of selected heroes, not players experimentally following the recommendations. It cannot be divided by the concordance gain from Part 11 to say how many times worse the product got.

Two panels. The left curve rises from split-half correlation 0.9261 at the binary target to 0.9374 near duration weight 0.75; its band shows standard deviation across half-splits. The right panel connects each of three seeds' validation top-one historical scores, all falling from about 56.30% to about 55.71% under duration weighting.
The repeatability check improved while the trained models' validation statistic fell. The left band is not a confidence interval; the right panel measures historical matchup outcomes, not wins caused by following advice.

A separate check found similar held-out win AUCs for simple pair estimates built from the two targets: approximately 0.5635 either way. That did not prove equal predictive value, and certainly didn't establish that the duration target was both strictly better and harmless.

I tried changing the loss to a least-squares form. The reported results remained close in those runs. Similar local behaviour of two losses can explain why that attempt buys little; it does not make the losses universally equivalent or the outcome inevitable before running them.

Finally, I kept the binary win objective and added a separate duration readout sharing the model's pair tables. Three tested weights again produced no useful improvement on the selected measures. Separating the output heads had not made the extra training signal helpful in this setup.

What the Gate Had Actually Passed

The gate had found more repeatable pair-effect estimates under a changed target. The training experiments asked whether that target improved the model on the metrics used to judge wins and candidate rankings. Those questions can have different answers.

I don't have evidence that the target contains a guaranteed benefit which this architecture simply cannot reach. It may emphasize the wrong distinctions, interact badly with the optimization, or help somewhere I did not test. The experiments close the tested approaches, not every possible use of duration.

The practical change is to make a cheap screening experiment exercise more of the intended path: target, training procedure, and downstream metric together. A repeatability diagnostic remains useful. It just doesn't get to stand in for that experiment.

The duration column stays in the dataset. It hasn't earned a place in this model's training objective.