This Saturday Hooker vs Parnasse · Accor Arena Paris France See this week's picks →
Methodology

How a fight probability is built, tested and graded.

This is the methods section: substrate, encoding, estimation, calibration, validation protocol and the machinery that makes every published pick auditable after the fact. It stops short of the parameterisation, which is not public. Everything else is here.

Live record bound at build time to the QC-signed canonical store · as of 2026-08-29 · 29 graded cards

Contents
  1. Data substrate
  2. Feature construction
  3. Estimator ensemble
  4. Calibration & scoring
  5. Validation protocol
  6. Provenance & public grading
  7. The predictability ceiling
  8. How to read a pick
  9. The empirical record
  10. Limitations & disclosures

In one paragraph. Every UFC bout is encoded as an orientation-invariant differential over point-in-time fighter state and, where a market exists, a de-vigged implied probability. An ensemble of independently trained estimators — gradient-boosted decision trees and regularised linear models, some anchored to the market price and some blind to it entirely — is combined by a logistic meta-learner fitted on out-of-sample outputs. The result is monotonically calibrated, scored under a proper scoring rule, validated walk-forward with a standing leakage battery, written once to a frozen artifact that is never rewritten, and graded in public afterwards whether it was right or wrong. §6 states exactly how much of that freeze we can prove and on how many cards.

How well that can possibly work is itself a question we researched rather than assumed. Our working paper, The Predictability Ceiling of UFC Fight Outcomes, measures the irreducible uncertainty of a UFC bout and finds the ceiling to be a property of the data rather than of any model — §7 states the finding and what follows from accepting it.

Across every pick we hold on file, 2023–2026 inclusive, that process has graded 1260-622 — 67% on n=1882. The 2026 live season, the only cohort published prospectively, stands at 245-105 (70%, n=350). Both are historical measurements, not forward promises.

01 · Substrate

Data substrate

The estimation set is the public professional record — bout outcomes and per-fight statistical lines reaching back to the early 1990s — joined to historical betting-market prices for the subset of bouts where a market existed. No private data, no paywalled feed, nothing a determined researcher could not assemble.

As-of construction, enforced by the pass order

The substrate is assembled in a single chronological pass. Every state variable a fighter carries into a bout — rating, form, activity, layoff, durability, strength of schedule — is advanced only after every bout on that date has been featurised. A fight's own result, and every result after it, is therefore structurally unavailable to the row that predicts it. This is an as-of(D) construction rather than a filtered one: the guarantee comes from the order of the pass, not from a date predicate somebody has to remember to write.

Market prices

Prices are taken from the latest pre-event snapshot strictly earlier than the event date, then converted to de-vigged (no-vig) implied probabilities — the bookmaker's margin removed so the two sides sum to one. That de-vigged probability is the market prior; it is never treated as ground truth.

Contamination controls

Career aggregates
Columns computed as of the scrape date — a fighter's lifetime accuracy as it stands today — are banned outright from the feature set. They are the classic backward leak: a 2019 row silently informed by 2026 results.
Regeneration control
Upstream regeneration can rewrite or duplicate settled bouts. Duplicate detection runs on the normalised fighter pair, and a bout appearing on two dates is quarantined rather than double-counted through a career, a rating and an evaluation window.
Survivorship
Full history is retained and down-weighted by an exponential half-life, never truncated to the era that flatters the model. Anomalous stretches — the empty-arena period among them — are weighted down, not excluded.
Coverage holes
Price coverage is incomplete in specific windows and on specific undercards. A bout with no reliable pre-event price is routed to the price-blind estimators rather than imputed, and the gap is disclosed instead of filled.
02 · Encoding

Feature construction

Encoding is symmetric by construction, because the raw record is not. The eventual winner is listed in the first corner far more often than chance would put them there — an artefact of how results are recorded, and a label sitting in the column order waiting to be learned.

Three properties remove it. Every feature is a differential — corner A minus corner B — plus context that is invariant to which corner is which. Every bout is presented to the estimators in both orientations during fitting. Predictions are averaged over the two orientations at inference. Corner order therefore carries no information any estimator can exploit, and the positional-inheritance class of defect — a value attached to the wrong fighter — is a build failure rather than a silent bias.

Feature families

Stated at category level. The feature set itself, its weighting and its parameterisation are not published.

Rating & uncertainty
Elo- and Glicko-family strength estimates with the rating deviation carried explicitly, so a debutant's uncertainty is modelled rather than hand-coded as a penalty.
Form & availability
Record, streak, layoff, age, and recency structure over recent bouts.
Output & efficiency
Recency-weighted per-minute rates: volume, accuracy, defensive rates, knockdown, takedown, submission and control activity.
Positional mix
Where the work happens — target distribution and position share — rather than raw totals.
Durability
Finish-loss history with recency weighting.
Schedule strength
Opponent quality, so a record built on a soft schedule is not read as a record built on a hard one.
Static attributes
Anthropometrics and stance — the only fields permitted to be time-invariant.
Market prior
The de-vigged implied probability. Supplied to the market-anchored estimators only; the price-blind path never sees it.
03 · Estimation

Estimator ensemble

The published probability is not one model's opinion. It is an ensemble of independently trained estimators — gradient-boosted decision trees and regularised linear models — combined by a logistic meta-learner: stacked generalization, with the meta-learner fitted on the members' out-of-sample outputs on a walk-forward schedule so it never sees a member's in-sample optimism.

Two functional classes

Market-anchored residual
The logit of the de-vigged market probability is supplied as the base margin. The estimator is shallow and heavily regularised, so it can only learn the deviations from the price that the data actually supports. A correction term on a strong prior — not a challenger to it.
Price-blind
No market information of any kind, at any stage. Whatever independent signal we hold lives in these estimators' disagreements with the price, which is exactly why they must never be allowed to see it.

What blending does and does not buy

Ensembling pays only to the extent members' errors are decorrelated, and on this problem they largely are not: measured error correlation across the members is high, because they read the same public record through similar lenses. The honest consequence is that combination buys robustness — insulation from any single estimator degrading, drifting or breaking — considerably more than it buys accuracy. We say that rather than sell the blend as the source of the number.

Bouts with no tradeable price — short-notice bookings, debut-heavy early prelims — are served by the price-blind path, so every fight on a card receives a probability. Full coverage is a product decision, not an accuracy claim: those bouts are genuinely harder to call, and the confidence tier reflects it rather than disguising it.

04 · Calibration

Calibration & scoring

A pick is a direction. A probability is a claim about frequency, and it is graded on a different axis.

Estimator outputs are mapped to probabilities by a monotone (isotonic-family) transform fitted strictly on past data. Monotone calibration is accuracy-invariant by construction — it cannot reorder picks, only restate their confidence — so a calibration change is a claim about the number, never about who we think wins. Fitting a calibrator on the full sample is a textbook subtle leak, invisible to any feature scan, and the walk-forward schedule below is what precludes it.

Proper scoring, not accuracy

Our primary metric is the Brier score — a strictly proper scoring rule, meaning it is minimised only by reporting your true belief, so it cannot be gamed by shading a number toward the side you want to look good on. Accuracy is reported because readers ask for it, and it is the second-order metric: it grades only the sign of the deviation from evenness, and is completely indifferent to whether a genuine coin flip was published at 52% or at 88%. A proper scoring rule is not indifferent, which is why it is the one we optimise and the one we lead with internally.

How a Brier is read

Never as one number. We take it in its calibration–refinement form — uncertainty − resolution + reliability, after Murphy and DeGroot. Uncertainty is the base-rate variance of the sport (0.245 on the historical cohort in §7); it is fixed, and no forecaster touches it. Resolution is discriminating power — how far a forecaster's conditional outcome rates depart from that base rate. Reliability is calibration error. A forecaster improves by driving reliability toward zero and pushing resolution as far as the available information allows, which is a shorter distance than the literature tends to assume — §7.

Primary metric · Brier0.2016
The mean squared error of the confidence percentage printed on the pick card, scored against what happened. Lower is better; a model that published 50% on everything would score 0.2500. This is a description of our own published probabilities and nothing else — no comparison to any market or benchmark is asserted here, and none belongs on this page until the serving-path question behind it is independently signed off. cohort: graded 2026 published picks · n=350 · recomputed from the receipts every build

Do not read that figure against a floor or a benchmark computed on some other cohort. A Brier score is only interpretable against its own sample: the base rate, the favourite–underdog mix and the price distribution all move it, so two cohorts a season apart are not the same test. The floors quoted in §7 belong to the cohorts they were estimated on, and are not a bar this number is being held against.

05 · Validation

Validation protocol

A model is only as honest as the test that produced its number. Five things keep ours straight, and each of them is run as a control that is expected to be able to fail.

  1. Walk-forward validation with strict temporal partitioning. Not a shuffled split. The estimators are refitted on a rolling schedule and every fit sees only bouts strictly earlier than its partition boundary — the way the model would actually have been used on the night. A random hold-out lets a model learn from fights that had not happened yet; it inflates every metric and means nothing.
  2. A standing leakage battery. Truncation invariance — features for a card must be identical when the substrate is physically truncated to that date, which tests the as-of guarantee rather than trusting it. Determinism — seeded re-runs must reproduce predictions exactly. Label-shuffle collapse — refitted on permuted outcomes, the price-blind estimators must fall to chance and the market-anchored ones must fall back to the market. A model that stays good on shuffled labels is reading something structural, and that is the single cheapest leak detector there is.
  3. Market-true nulls. Where a claim is about beating a benchmark, the null distribution is resampled against the real price distribution, not a naive label shuffle. Shuffle nulls are far too easy to clear and will certify almost anything; several results that looked significant against a shuffle died against a market-true null, which is the correct outcome.
  4. Pre-registered evaluation gates. The primary metric, its threshold, the tie-break and the persistence requirement are filed before the evaluation runs. Choosing the bar afterwards is how "it beat the baseline" quietly becomes metric shopping — accuracy says no, macro-F1 says yes, publish the one that says yes. A result that clears a bar written after the fact is not a result.
  5. Adversarial external rebuilds. Independent builds from the same public substrate — some run genuinely blind to our architecture and our figures — are commissioned specifically to break the number. They have found real defects, which is the point of paying for them, and they have repeatedly landed in the same accuracy band. See Limitations: we read that as evidence about the ceiling of this data, not as a compliment.
06 · Provenance

Provenance & public grading

A forecast that can be edited after the event is not a forecast. Most of the engineering effort here is not in the model at all — it is in making the prediction impossible to quietly revise, and the grading impossible to quietly flatter.

Write-once freeze
Each card's prediction set is written once to a separate frozen artifact and that copy is never rewritten: the live export may be regenerated, the witness may not. What that proves depends on when the witness was written, so we count it rather than assert it. Of the 350 graded 2026 picks, 31 — 3 cards of 29 (2026-05-30, 2026-06-07, 2026-06-15) — carry a genuine pre-card export timestamp on the frozen copy. On the remaining 319 the frozen_at field is a placeholder and the export was regenerated after the bout, so those picks are write-once but not provably pre-bout from the artifact itself. We hold no evidence they were altered; we also cannot show you that they weren't, and the difference is the whole point of this section.
Content-hash provenance
The record store this site's every published figure is bound to — the canonical season record — is SHA-256-checked against the fingerprint its publisher issued, before a page is written, and a mismatch aborts the build. Hash, never modification time: every proxy check we have relied on has eventually been fooled by something that looked fresh and was not. A content hash cannot be. The other artifacts we derive on the way to a page carry a recorded hash of their own output and of every input, which makes a disagreement diagnosable — that is a weaker guarantee than the one above and it is not the one holding up the numbers.
Entity-resolved grading
Predictions join results on the normalised fighter pair — accent- and punctuation-folded, order-independent — never on row position and never on a regenerable identifier. Positional and id joins are not a hypothetical failure here: one produced a mis-graded card that reconciled to a number present in no source file.
Independent reconciliation
Rendered per-card tallies are asserted against the canonical record for every graded card, and any disagreement fails the build rather than publishing. Freshness is checked separately from correctness, because a number that is current and wrong passes a freshness check.
Adversarial re-grading
An independent lane re-derives published figures from the raw artifacts with an explicit mandate to falsify them. Figures found unreproducible are withdrawn, not defended — the provisional flag under Limitations is one such finding, still standing against us.
Quarantine
A card whose export cannot be reconstructed is withheld from the public table and counted as withheld, so the row accounting still adds up. Silently dropping an inconvenient card is the one failure mode a public record cannot survive.

tier provenance · write-once tier freeze · 28 card(s) provisional · 0 with post-freeze drift

07 · The wall

The predictability ceiling

Published UFC winner models cluster in the 65–70% band whatever the architecture — logistic regression, boosted ensemble or neural network, market-aware or market-blind. We did not treat that as a plateau waiting for a better feature. We treated the convergence itself as the object of study, and measured it: The Predictability Ceiling of UFC Fight Outcomes, our working paper.

The question is not how do we beat 70%? but why 70? — because the two answers imply opposite research programmes. If the band is a modelling limit, the response is better features and architectures. If it is an aleatoric limit — set by the irreducible randomness of a single fight, where one landed strike ends a bout regardless of who was the better fighter — then further effort spent on winner accuracy is misdirected, and the honest contribution is to characterise the bound.

Three findings

  1. There is a floor, and the price is already close to it. Decomposing outcome uncertainty over 2,502 historical bouts with de-vigged closing prices, the isotonically recalibrated aleatoric floor — the best Brier score attainable by any forecaster sharing that information — sits at a Brier of 0.205, against 0.245 for a forecaster who knows only the base rate. Nearly all of the distance between knowing nothing and knowing everything knowable has been travelled by the time you have the closing price.
  2. Most of a fight is noise, and the share is measurable. In information terms a bout outcome carries 0.98 bits of entropy, of which the closing market extracts 0.11 bits — leaving roughly 89% irreducible. Given two fighters and everything knowable before the walk, the winner remains a high-variance coin weighted only slightly by skill.
  3. The ceiling is in the data, not the model. Across four independent predictors — the market, an odds-aware model, an odds-blind model and their blend — none out-resolves the market. Were 70% a modelling limit, something would. That is the signature of a bound living in the information available about a fight rather than in the function class used to fit it, and it reproduces out-of-sample on a pre-registered frozen forward season (n=253 across 21 cards, weights frozen before the season opened): the model reaches the wall and does not cross it. That forward season has since run on: the same pre-registered cohort now stands at 350 graded fights across 29 cards (70%), and the extended sample sits in the same band. The wall did not move when the sample grew, which is what a real bound does and what a lucky one does not.

Two things follow, and both are already on this page. Calibration over accuracy — when accuracy is bounded and nearly saturated it is a weak discriminator between models, and reliability and resolution are the informative axes, which is why §4 leads on a proper scoring rule. And a standing prior: a model reporting materially higher walk-forward accuracy on real fights is extraordinary, and should be treated as a leak until proven otherwise — including when it is ours. That prior is what the battery in §5 exists to serve.

The same prior applies to short-window betting results, and the literature has a clean cautionary case: the one prominent live-betting result in this field was withdrawn by its own author in 2026, after a longer sample showed substantially weaker performance than the eight-week window it was announced on. A short window against a single book is the highest-variance, highest-mirage quantity in this domain. It is exactly what a bound-limited process produces when you look at it briefly, and exactly why nothing on this site is presented as a betting result.

What has changed since the paper

The paper's out-of-sample protocol — freeze the weights, publish before the card, never tune on the forward sample — was a promise kept by discipline. It is now kept by mechanism. Pre-event artifacts are write-once and witnessed by an externally timestamped copy we do not control; every number rendered on this site asserts a content hash against a canonical manifest before the page is written; and grading is entity-keyed and re-derived adversarially from the raw artifacts rather than trusted. That machinery is §6, and it is the substantive difference between a protocol described in a paper and a protocol a reader can check.

The constructive corollary — confidence with a coverage guarantee

If the point prediction is capped, the quantity worth optimising is the honesty of the confidence attached to it. The paper's corollary shows that split-conformal prediction turns a fight model's raw probabilities into nested confidence sets carrying a distribution-free, finite-sample coverage guarantee: thresholds fitted on an earlier block of the frozen season (n=151) and evaluated on a later one (n=102), split chronologically at a card boundary so no future fight can inform a past threshold. Realised accuracy landed inside its target interval at every level, and — unlike a conventional fixed-threshold tiering fitted on the same data — the tiers came out monotone out-of-sample. Each step up the ladder really was a step up.

The confidence ladder in the next section is the product expression of that corollary — the method, applied, with a fifth step the paper's four-set construction does not have: an explicit TOSS-UP floor for bouts no confidence set can separate. A tier is meant to be a confidence set with a number we are held to, not a label, which is the whole reason this site publishes per-tier hit rates instead of a single headline. The rates you see there are the live ones, recomputed from the graded record every build — not the paper's held-out figures, which belong to its own cohort. The paper's two caveats travel with the method regardless: coverage is marginal over the score distribution rather than conditional on weight class or favourite status, and the top tiers are thinly populated, so their floors are not yet provably tight at the sample sizes involved. The method is sound; the high tiers need more fights. The live numbers below are how you watch that happen instead of taking it on faith.

Finally, the bound is conditional on public pre-fight information, and says nothing beyond it. Two directions escape it in principle — signals genuinely orthogonal to the closing price, and prediction during the fight itself, where the relevant market is a different one. Both are open problems. Neither is solved by a better pre-fight winner model, which is precisely the point.

source: our working paper, The Predictability Ceiling of UFC Fight Outcomes · calibration–refinement (Murphy/DeGroot) decomposition, information accounting in bits, split-conformal corollary, nonparametric bootstrap intervals · figures above are the paper's own cohorts, are not comparable with the live-season figures elsewhere on this page, and are superseded for the current season by the public record — which has since grown and been re-graded. The record page is authoritative.

08 · Reading the output

How to read a pick — the confidence tiers

The rest of this page is in plain English, because a well-calibrated probability that nobody reads correctly is a wasted probability. Every pick lands in one of five tiers by how confident the ensemble is. Higher tier, higher conviction — that is what a tier is: a statement about how tightly the estimators agree, not a promise about the hit rate. LOCK (80%, 20-5) leads the live ladder, but the rungs below it are not a clean staircase: MED (79%, 79-21) currently out-hits HIGH (78.26%, 54-15). At these sample sizes that is what to expect, and we would rather say so than round it into a story: the 95% Wilson intervals of every neighbouring pair overlap, so no two adjacent tiers are statistically distinguishable yet. Read the gaps between neighbours as noise, not as a ranking. What the season does separate is the two ends of the ladder: LOCK (80%, 20-5) against TOSS-UP (49.37%, 39-40) is a real gap, and it is the gap the tier is for. These are the real live numbers (2026-08-29, 29 cards), not a target.

LOCK Highest conviction
80%
20-5 · n=25
HIGH Strong read
78.26%
54-15 · n=69
MED Solid lean
79%
79-21 · n=100
LOW Slight edge
68.83%
53-24 · n=77
TOSS-UP A coin flip — and we say so
49.37%
39-40 · n=79

Read it top to bottom: our LOCK and HIGH calls are where we have the most conviction. A TOSS-UP is exactly that — a coin flip, and we label it one. We never say to bet a toss-up; it is shown for honesty, and it stays in the record.

The percentage and the tier are two different things

A pick carries both a number and a tier, and they can disagree — so it is worth being explicit that the tier is not simply a band cut out of the percentage. The percentage is the ensemble's point estimate that this fighter wins. The tier is how much conviction that estimate carries, and it also reflects how tightly the independent estimators agree with each other and where the conformal thresholds of §7 fall.

The practical consequence, which surprises people: a pick shown in the mid-70s that the estimators disagree about can sit a tier below a pick in the mid-60s they are unanimous on. That is not a display bug and it is not a typo — it is the ladder doing its job, marking down a confident-looking number that only one read supports. The percentage is the estimate; the tier is the confidence attached to it. If you only look at one of them, look at the tier.

01

Every fight, modelled

Each bout runs through the whole ensemble — several independent reads of the same fight, built on the full public record and the live market.

02

Blended into one read

The meta-learner combines them into a single probability. Where the estimators agree, confidence rises; where they disagree, it falls — honestly.

03

One confidence %

You get one pick per fight, one calibrated number and one tier, from a coin flip up to a lock. The number is the estimate; the tier is the conviction behind it.

04

Graded in public

The pick is frozen before the cage door shuts and graded after. Wins, losses and toss-ups all go on the public record.

09 · Results

The empirical record

The headline is the every-fight number, because it is the whole truth: every fight we called in 2026, toss-ups included, nothing dropped.

70%
245-105 — every fight we called in 2026 (n=350).
every pick · incl. toss-ups
Why we grade all of them. A tipster who only shows their winners can look unbeatable. Selective disclosure is not a small sin in forecasting — it is the whole difference between a track record and a highlight reel. We show every call, including the coin flips and the ones we got wrong, because the honest full-season number is the only one worth trusting.

The confidence gradient

Not every call is equal. When the independent estimators line up on the same fighter, the pick sharpens — and the record shows it.

77.78%
63-18 when the estimators strongly agree (n=81).
strong consensus
74.7%
124-42 whenever the estimators agree at all (n=166).
estimators agree
67%
1260-622 across every graded pick on file, 2023–2026 (n=1882).
all cohorts · incl. backtest

Those cohorts answer different questions and must never be blended into a single figure. The 2026 season is the only one published prospectively — frozen before the card, graded after. Earlier seasons are reconstructions of what the pipeline would have said, and are labelled as such wherever they appear.

10 · Limitations

Limitations & disclosures

The parts a methods section is obliged to state plainly. Every number on this page carries its sample size and its cohort; here is what they do — and do not — mean.

See this week's picks → Every graded pick