Held out from the final model-selection decision.
MODEL EVIDENCE · HISTORICAL + CURRENT · ACCURACY
Accuracy
INDEPENDENT PROMOTION GATE
Does the model actually work?
YES — it passed the independent promotion gate.
The final game-win model was chosen on earlier seasons and then judged on 5,282 games it did not use to choose the champion. It beat Classic Elo on the probabilistic scoring rules used for promotion.
Out of 5,282 formal OOS games.
Thresholded at 50%. Not the only promotion metric.
39 more correct winner calls across the same 5,282 games.
2026–27 TEMPORAL TRACKING
Accuracy by forecast horizon
Live 2026–27 accuracy is not available until verified final games exist. The tracking contract is active before opening night; no preseason hit rate is fabricated.
| Checkpoint | Games | Coverage | Winner accuracy | Log Loss | Brier | Calibration ECE |
|---|---|---|---|---|---|---|
| Early forecast | 0 | — | — | — | — | — |
| 7-day | 0 | — | — | — | — | — |
| 24-hour | 0 | — | — | — | — | — |
| Final pregame | 0 | — | — | — | — | — |
Pregame only. Release first_seen_at_utc controls eligibility. Missing checkpoints are withheld. Log Loss, Brier, winner accuracy and calibration use verified final results only.
PROBHOOPS VS CLASSIC ELO
Why the champion was promoted
Winner accuracy is easy to understand, but probability quality is the stricter test. Lower Log Loss and Brier mean the probabilities themselves were better, not merely the final side of 50%.
Adjusted Efficiency + Elo Bridge
- Winner accuracy
- 65.66%
- Correct calls
- 3,468
- Log Loss
- 0.618471
- Brier
- 0.214950
- Calibration ECE
- 1.83%
Classic Elo
- Winner accuracy
- 64.92%
- Correct calls
- 3,429
- Log Loss
- 0.624031
- Brier
- 0.217445
- Calibration ECE
- 1.89%
TEMPORAL FOLDS
Did the edge survive year by year?
Yes on the primary probability score: ProbHoops recorded lower Log Loss than Classic Elo in every held-out season. Accuracy itself can move differently in an individual season, which is why it is not used alone.
UNCERTAINTY CHECK
The improvement survives clustered resampling
95% CI -0.008954 to -0.002162. The entire interval is below zero.
5,000 bootstrap draws · 846 NBA-date clusters95% CI -0.003965 to -0.001005. The entire interval is below zero.
5,000 bootstrap draws · 846 NBA-date clustersSCORE PROJECTION
What does the score model miss by?
Average absolute error for one team’s final score.
Form baseline: ±9.57Average absolute error in the final point differential.
Form baseline: ±11.34Average absolute error in combined points.
Form baseline: ±15.27The promoted score model beat the form-score baseline on all three error measures in every formal OOS season.
CALIBRATION
When ProbHoops says X%, what happened?
A calibrated model should see outcomes occur at roughly the rate it announces. The bars below compare predicted home-win probability with the observed result frequency.
PLAYER-AWARE CHALLENGER · SHADOW ONLY
A second model passed its research gates — without replacing the champion
Adding pregame player-role fragility improved the historical challenger from 65.66% to 66.15%, or +26 correct winners. The accepted public champion is still unchanged because operational current-player availability and expected minutes require a separate gate.
3,494 correct of 5,282.
Incumbent 0.618471.
Incumbent 0.214950.
Evidence supported; public probability activation remains off.
Bootstrap Log Loss Δ: -0.002580 with 95% CI -0.004657 to -0.000451. The interval remains below zero.
PLAYER MILESTONES · SEPARATE MODEL
Conditional-on-active calibration
Strong historical signal — different question.
84.45% threshold accuracy belongs to the Player Milestone shadow model. It means Conditional on player being active and is not game-win accuracy. It must never be compared directly with the 65.66% game-winner figure.
Historical shadow validation only.
Conditional-on-active probability quality.
Lower is better.
Formal OOS 2022–2025.
TECHNICAL LAYER
What the metrics mean
Strongly penalizes confident probabilities that are wrong. Lower is better.
Mean squared error of predicted probabilities. Lower is better.
Measures how closely announced probabilities match observed frequencies. Lower is better.
