PROBHOOPS · NBA ANALYTICS · TEAMS · PLAYERS · PROBABILITIES
FIND ANYTHING

Teams & players

Start typing to search the ProbHoops directory.

Tip: press / anywhere to search.

MODEL EVIDENCE · HISTORICAL + CURRENT · ACCURACY

Accuracy

INDEPENDENT PROMOTION GATE

Does the model actually work?

YES — it passed the independent promotion gate.

The final game-win model was chosen on earlier seasons and then judged on 5,282 games it did not use to choose the champion. It beat Classic Elo on the probabilistic scoring rules used for promotion.

FORMAL OOSPASS2022–2025
UNSEEN GAMES5,282

Held out from the final model-selection decision.

WINNERS CALLED CORRECTLY3,468

Out of 5,282 formal OOS games.

WINNER ACCURACY65.66%

Thresholded at 50%. Not the only promotion metric.

VS CLASSIC ELO+39

39 more correct winner calls across the same 5,282 games.

2026–27 TEMPORAL TRACKING

Accuracy by forecast horizon

preseason ready

Live 2026–27 accuracy is not available until verified final games exist. The tracking contract is active before opening night; no preseason hit rate is fabricated.

CheckpointGamesCoverageWinner accuracyLog LossBrierCalibration ECE
Early forecast0—————
7-day0—————
24-hour0—————
Final pregame0—————
Early first eligible release containing the game before tipoff7-day latest game-containing release at or before T−7d24-hour latest game-containing release at or before T−24hFinal pregame latest game-containing release before tipoff

Pregame only. Release first_seen_at_utc controls eligibility. Missing checkpoints are withheld. Log Loss, Brier, winner accuracy and calibration use verified final results only.

PROBHOOPS VS CLASSIC ELO

Why the champion was promoted

4/4 OOS seasons · Log Loss

Winner accuracy is easy to understand, but probability quality is the stricter test. Lower Log Loss and Brier mean the probabilities themselves were better, not merely the final side of 50%.

BENCHMARKCLASSIC ELO

Classic Elo

Winner accuracy
64.92%
Correct calls
3,429
Log Loss
0.624031
Brier
0.217445
Calibration ECE
1.89%
ACCURACY+0.74 ppProbHoops − Elo
LOG LOSS-0.005560negative is better
BRIER-0.002494negative is better
SEASON STABILITY4/4lower Log Loss than Elo

TEMPORAL FOLDS

Did the edge survive year by year?

no cherry-picked season

Yes on the primary probability score: ProbHoops recorded lower Log Loss than Classic Elo in every held-out season. Accuracy itself can move differently in an individual season, which is why it is not used alone.

2022–231,320 games
ProbHoops accuracy63.0%
Classic Elo accuracy62.5%
Log Loss · ProbHoops0.6472
Log Loss · Elo0.6499
✓ ProbHoops lower Log Loss
2023–241,319 games
ProbHoops accuracy65.5%
Classic Elo accuracy64.4%
Log Loss · ProbHoops0.6139
Log Loss · Elo0.6152
✓ ProbHoops lower Log Loss
2024–251,321 games
ProbHoops accuracy65.3%
Classic Elo accuracy65.6%
Log Loss · ProbHoops0.6117
Log Loss · Elo0.6187
✓ ProbHoops lower Log Loss
2025–261,322 games
ProbHoops accuracy68.8%
Classic Elo accuracy67.2%
Log Loss · ProbHoops0.6012
Log Loss · Elo0.6123
✓ ProbHoops lower Log Loss

UNCERTAINTY CHECK

The improvement survives clustered resampling

NBA date bootstrap
LOG LOSS Δ · PROBHOOPS − ELO-0.005560

95% CI -0.008954 to -0.002162. The entire interval is below zero.

5,000 bootstrap draws · 846 NBA-date clusters
BRIER Δ · PROBHOOPS − ELO-0.002494

95% CI -0.003965 to -0.001005. The entire interval is below zero.

5,000 bootstrap draws · 846 NBA-date clusters

SCORE PROJECTION

What does the score model miss by?

5,282 OOS games
TEAM SCORE ERROR±9.43 pts

Average absolute error for one team’s final score.

Form baseline: ±9.57
MARGIN ERROR±11.02 pts

Average absolute error in the final point differential.

Form baseline: ±11.34
GAME TOTAL ERROR±15.12 pts

Average absolute error in combined points.

Form baseline: ±15.27

The promoted score model beat the form-score baseline on all three error measures in every formal OOS season.

CALIBRATION

When ProbHoops says X%, what happened?

ECE 1.83%

A calibrated model should see outcomes occur at roughly the rate it announces. The bars below compare predicted home-win probability with the observed result frequency.

10–20%50 games
Predicted16.1%
Observed16.0%
20–30%263 games
Predicted25.9%
Observed17.1%
30–40%542 games
Predicted35.3%
Observed31.4%
40–50%847 games
Predicted45.4%
Observed43.2%
50–60%1,270 games
Predicted55.4%
Observed55.6%
60–70%1,121 games
Predicted64.8%
Observed63.2%
70–80%777 games
Predicted74.7%
Observed75.9%
80–90%372 games
Predicted83.9%
Observed84.1%
90–100%39 games
Predicted91.7%
Observed94.9%

PLAYER-AWARE CHALLENGER · SHADOW ONLY

A second model passed its research gates — without replacing the champion

5,282 same OOS games

Adding pregame player-role fragility improved the historical challenger from 65.66% to 66.15%, or +26 correct winners. The accepted public champion is still unchanged because operational current-player availability and expected minutes require a separate gate.

WINNER ACCURACY66.15%

3,494 correct of 5,282.

LOG LOSS0.615891

Incumbent 0.618471.

BRIER0.213745

Incumbent 0.214950.

STATUSSHADOW

Evidence supported; public probability activation remains off.

Bootstrap Log Loss Δ: -0.002580 with 95% CI -0.004657 to -0.000451. The interval remains below zero.

PLAYER MILESTONES · SEPARATE MODEL

Conditional-on-active calibration

Strong historical signal — different question.

1,527,999 OOS events

84.45% threshold accuracy belongs to the Player Milestone shadow model. It means Conditional on player being active and is not game-win accuracy. It must never be compared directly with the 65.66% game-winner figure.

THRESHOLD ACCURACY84.45%

Historical shadow validation only.

LOG LOSS0.345393

Conditional-on-active probability quality.

BRIER0.109081

Lower is better.

CALIBRATION ECE0.32%

Formal OOS 2022–2025.

MODEL ≥80%88.89% observed144,477 OOS events · mean predicted 89.49%
MODEL ≥90%93.54% observed71,138 OOS events · mean predicted 93.93%
MODEL ≥95%96.28% observed24,489 OOS events · mean predicted 95.79%

TECHNICAL LAYER

What the metrics mean

transparent by design
LOG LOSSConfidence matters

Strongly penalizes confident probabilities that are wrong. Lower is better.

BRIERProbability error

Mean squared error of predicted probabilities. Lower is better.

ECECalibration gap

Measures how closely announced probabilities match observed frequencies. Lower is better.