Every prediction, every event, tracked and published — here's how the model has actually performed.
Note: Accuracy statistics reflect winner predictions only. Method of victory, round predictions, and finish times are provided as supplemental analysis but are not included in these performance metrics.
Not all predictions are equal. Our model assigns a confidence level (High/Medium/Low) to every pick before the fight happens, and the High tier is where it has been clearly strongest. Here's how each tier has actually performed:
Anyone can publish an accuracy number. The harder question: when we say a fight is 70%, does it actually win about 70% of the time? That's calibration, and it's the real test of whether a confidence number means anything.
Each dot is a confidence band. The dashed line is perfect calibration — a dot sitting on it means we won exactly as often as we said we would.
| Confidence Band | Fights | We Said | Actually Won | Gap |
|---|---|---|---|---|
| 50-54% | 79 | 52.2% | 50.6% | -1.6 |
| 55-59% | 89 | 56.8% | 64.0% | +7.3 |
| 60-64% | 72 | 61.8% | 56.9% | -4.8 |
| 65-69% | 36 | 66.7% | 66.7% | -0.0 |
| 70-74% | 41 | 71.7% | 68.3% | -3.4 |
| 75-79% | 31 | 77.1% | 80.6% | +3.5 |
| 80-84% | 19 | 81.8% | 94.7% | +12.9 |
| 90-94% | 22 | 90.0% | 81.8% | -8.2 |
Bands with fewer than 10 graded fights are left off — not enough data yet to mean anything. Red means we claimed more confidence than we delivered. Green means we delivered more than we claimed — a good surprise, not a problem.