This is the methods section: substrate, encoding, estimation, calibration, validation protocol and the machinery that makes every published pick auditable after the fact. It stops short of the parameterisation, which is not public. Everything else is here.
Live record bound at build time to the QC-signed canonical store · as of 2026-08-29 · 29 graded cards
In one paragraph. Every UFC bout is encoded as an orientation-invariant differential over point-in-time fighter state and, where a market exists, a de-vigged implied probability. An ensemble of independently trained estimators — gradient-boosted decision trees and regularised linear models, some anchored to the market price and some blind to it entirely — is combined by a logistic meta-learner fitted on out-of-sample outputs. The result is monotonically calibrated, scored under a proper scoring rule, validated walk-forward with a standing leakage battery, written once to a frozen artifact that is never rewritten, and graded in public afterwards whether it was right or wrong. §6 states exactly how much of that freeze we can prove and on how many cards.
How well that can possibly work is itself a question we researched rather than assumed. Our working paper, The Predictability Ceiling of UFC Fight Outcomes, measures the irreducible uncertainty of a UFC bout and finds the ceiling to be a property of the data rather than of any model — §7 states the finding and what follows from accepting it.
Across every pick we hold on file, 2023–2026 inclusive, that process has graded 1260-622 — 67% on n=1882. The 2026 live season, the only cohort published prospectively, stands at 245-105 (70%, n=350). Both are historical measurements, not forward promises.
The estimation set is the public professional record — bout outcomes and per-fight statistical lines reaching back to the early 1990s — joined to historical betting-market prices for the subset of bouts where a market existed. No private data, no paywalled feed, nothing a determined researcher could not assemble.
The substrate is assembled in a single chronological pass. Every state variable a fighter carries into a bout — rating, form, activity, layoff, durability, strength of schedule — is advanced only after every bout on that date has been featurised. A fight's own result, and every result after it, is therefore structurally unavailable to the row that predicts it. This is an as-of(D) construction rather than a filtered one: the guarantee comes from the order of the pass, not from a date predicate somebody has to remember to write.
Prices are taken from the latest pre-event snapshot strictly earlier than the event date, then converted to de-vigged (no-vig) implied probabilities — the bookmaker's margin removed so the two sides sum to one. That de-vigged probability is the market prior; it is never treated as ground truth.
Encoding is symmetric by construction, because the raw record is not. The eventual winner is listed in the first corner far more often than chance would put them there — an artefact of how results are recorded, and a label sitting in the column order waiting to be learned.
Three properties remove it. Every feature is a differential — corner A minus corner B — plus context that is invariant to which corner is which. Every bout is presented to the estimators in both orientations during fitting. Predictions are averaged over the two orientations at inference. Corner order therefore carries no information any estimator can exploit, and the positional-inheritance class of defect — a value attached to the wrong fighter — is a build failure rather than a silent bias.
Stated at category level. The feature set itself, its weighting and its parameterisation are not published.
The published probability is not one model's opinion. It is an ensemble of independently trained estimators — gradient-boosted decision trees and regularised linear models — combined by a logistic meta-learner: stacked generalization, with the meta-learner fitted on the members' out-of-sample outputs on a walk-forward schedule so it never sees a member's in-sample optimism.
Ensembling pays only to the extent members' errors are decorrelated, and on this problem they largely are not: measured error correlation across the members is high, because they read the same public record through similar lenses. The honest consequence is that combination buys robustness — insulation from any single estimator degrading, drifting or breaking — considerably more than it buys accuracy. We say that rather than sell the blend as the source of the number.
Bouts with no tradeable price — short-notice bookings, debut-heavy early prelims — are served by the price-blind path, so every fight on a card receives a probability. Full coverage is a product decision, not an accuracy claim: those bouts are genuinely harder to call, and the confidence tier reflects it rather than disguising it.
A pick is a direction. A probability is a claim about frequency, and it is graded on a different axis.
Estimator outputs are mapped to probabilities by a monotone (isotonic-family) transform fitted strictly on past data. Monotone calibration is accuracy-invariant by construction — it cannot reorder picks, only restate their confidence — so a calibration change is a claim about the number, never about who we think wins. Fitting a calibrator on the full sample is a textbook subtle leak, invisible to any feature scan, and the walk-forward schedule below is what precludes it.
Our primary metric is the Brier score — a strictly proper scoring rule, meaning it is minimised only by reporting your true belief, so it cannot be gamed by shading a number toward the side you want to look good on. Accuracy is reported because readers ask for it, and it is the second-order metric: it grades only the sign of the deviation from evenness, and is completely indifferent to whether a genuine coin flip was published at 52% or at 88%. A proper scoring rule is not indifferent, which is why it is the one we optimise and the one we lead with internally.
Never as one number. We take it in its calibration–refinement form — uncertainty − resolution + reliability, after Murphy and DeGroot. Uncertainty is the base-rate variance of the sport (0.245 on the historical cohort in §7); it is fixed, and no forecaster touches it. Resolution is discriminating power — how far a forecaster's conditional outcome rates depart from that base rate. Reliability is calibration error. A forecaster improves by driving reliability toward zero and pushing resolution as far as the available information allows, which is a shorter distance than the literature tends to assume — §7.
Do not read that figure against a floor or a benchmark computed on some other cohort. A Brier score is only interpretable against its own sample: the base rate, the favourite–underdog mix and the price distribution all move it, so two cohorts a season apart are not the same test. The floors quoted in §7 belong to the cohorts they were estimated on, and are not a bar this number is being held against.
A model is only as honest as the test that produced its number. Five things keep ours straight, and each of them is run as a control that is expected to be able to fail.
A forecast that can be edited after the event is not a forecast. Most of the engineering effort here is not in the model at all — it is in making the prediction impossible to quietly revise, and the grading impossible to quietly flatter.
tier provenance · write-once tier freeze · 28 card(s) provisional · 0 with post-freeze drift
Published UFC winner models cluster in the 65–70% band whatever the architecture — logistic regression, boosted ensemble or neural network, market-aware or market-blind. We did not treat that as a plateau waiting for a better feature. We treated the convergence itself as the object of study, and measured it: The Predictability Ceiling of UFC Fight Outcomes, our working paper.
The question is not how do we beat 70%? but why 70? — because the two answers imply opposite research programmes. If the band is a modelling limit, the response is better features and architectures. If it is an aleatoric limit — set by the irreducible randomness of a single fight, where one landed strike ends a bout regardless of who was the better fighter — then further effort spent on winner accuracy is misdirected, and the honest contribution is to characterise the bound.
Two things follow, and both are already on this page. Calibration over accuracy — when accuracy is bounded and nearly saturated it is a weak discriminator between models, and reliability and resolution are the informative axes, which is why §4 leads on a proper scoring rule. And a standing prior: a model reporting materially higher walk-forward accuracy on real fights is extraordinary, and should be treated as a leak until proven otherwise — including when it is ours. That prior is what the battery in §5 exists to serve.
The same prior applies to short-window betting results, and the literature has a clean cautionary case: the one prominent live-betting result in this field was withdrawn by its own author in 2026, after a longer sample showed substantially weaker performance than the eight-week window it was announced on. A short window against a single book is the highest-variance, highest-mirage quantity in this domain. It is exactly what a bound-limited process produces when you look at it briefly, and exactly why nothing on this site is presented as a betting result.
The paper's out-of-sample protocol — freeze the weights, publish before the card, never tune on the forward sample — was a promise kept by discipline. It is now kept by mechanism. Pre-event artifacts are write-once and witnessed by an externally timestamped copy we do not control; every number rendered on this site asserts a content hash against a canonical manifest before the page is written; and grading is entity-keyed and re-derived adversarially from the raw artifacts rather than trusted. That machinery is §6, and it is the substantive difference between a protocol described in a paper and a protocol a reader can check.
If the point prediction is capped, the quantity worth optimising is the honesty of the confidence attached to it. The paper's corollary shows that split-conformal prediction turns a fight model's raw probabilities into nested confidence sets carrying a distribution-free, finite-sample coverage guarantee: thresholds fitted on an earlier block of the frozen season (n=151) and evaluated on a later one (n=102), split chronologically at a card boundary so no future fight can inform a past threshold. Realised accuracy landed inside its target interval at every level, and — unlike a conventional fixed-threshold tiering fitted on the same data — the tiers came out monotone out-of-sample. Each step up the ladder really was a step up.
The confidence ladder in the next section is the product expression of that corollary — the method, applied, with a fifth step the paper's four-set construction does not have: an explicit TOSS-UP floor for bouts no confidence set can separate. A tier is meant to be a confidence set with a number we are held to, not a label, which is the whole reason this site publishes per-tier hit rates instead of a single headline. The rates you see there are the live ones, recomputed from the graded record every build — not the paper's held-out figures, which belong to its own cohort. The paper's two caveats travel with the method regardless: coverage is marginal over the score distribution rather than conditional on weight class or favourite status, and the top tiers are thinly populated, so their floors are not yet provably tight at the sample sizes involved. The method is sound; the high tiers need more fights. The live numbers below are how you watch that happen instead of taking it on faith.
Finally, the bound is conditional on public pre-fight information, and says nothing beyond it. Two directions escape it in principle — signals genuinely orthogonal to the closing price, and prediction during the fight itself, where the relevant market is a different one. Both are open problems. Neither is solved by a better pre-fight winner model, which is precisely the point.
source: our working paper, The Predictability Ceiling of UFC Fight Outcomes · calibration–refinement (Murphy/DeGroot) decomposition, information accounting in bits, split-conformal corollary, nonparametric bootstrap intervals · figures above are the paper's own cohorts, are not comparable with the live-season figures elsewhere on this page, and are superseded for the current season by the public record — which has since grown and been re-graded. The record page is authoritative.
The rest of this page is in plain English, because a well-calibrated probability that nobody reads correctly is a wasted probability. Every pick lands in one of five tiers by how confident the ensemble is. Higher tier, higher conviction — that is what a tier is: a statement about how tightly the estimators agree, not a promise about the hit rate. LOCK (80%, 20-5) leads the live ladder, but the rungs below it are not a clean staircase: MED (79%, 79-21) currently out-hits HIGH (78.26%, 54-15). At these sample sizes that is what to expect, and we would rather say so than round it into a story: the 95% Wilson intervals of every neighbouring pair overlap, so no two adjacent tiers are statistically distinguishable yet. Read the gaps between neighbours as noise, not as a ranking. What the season does separate is the two ends of the ladder: LOCK (80%, 20-5) against TOSS-UP (49.37%, 39-40) is a real gap, and it is the gap the tier is for. These are the real live numbers (2026-08-29, 29 cards), not a target.
Read it top to bottom: our LOCK and HIGH calls are where we have the most conviction. A TOSS-UP is exactly that — a coin flip, and we label it one. We never say to bet a toss-up; it is shown for honesty, and it stays in the record.
A pick carries both a number and a tier, and they can disagree — so it is worth being explicit that the tier is not simply a band cut out of the percentage. The percentage is the ensemble's point estimate that this fighter wins. The tier is how much conviction that estimate carries, and it also reflects how tightly the independent estimators agree with each other and where the conformal thresholds of §7 fall.
The practical consequence, which surprises people: a pick shown in the mid-70s that the estimators disagree about can sit a tier below a pick in the mid-60s they are unanimous on. That is not a display bug and it is not a typo — it is the ladder doing its job, marking down a confident-looking number that only one read supports. The percentage is the estimate; the tier is the confidence attached to it. If you only look at one of them, look at the tier.
Each bout runs through the whole ensemble — several independent reads of the same fight, built on the full public record and the live market.
The meta-learner combines them into a single probability. Where the estimators agree, confidence rises; where they disagree, it falls — honestly.
You get one pick per fight, one calibrated number and one tier, from a coin flip up to a lock. The number is the estimate; the tier is the conviction behind it.
The pick is frozen before the cage door shuts and graded after. Wins, losses and toss-ups all go on the public record.
The headline is the every-fight number, because it is the whole truth: every fight we called in 2026, toss-ups included, nothing dropped.
Not every call is equal. When the independent estimators line up on the same fighter, the pick sharpens — and the record shows it.
Those cohorts answer different questions and must never be blended into a single figure. The 2026 season is the only one published prospectively — frozen before the card, graded after. Earlier seasons are reconstructions of what the pipeline would have said, and are labelled as such wherever they appear.
The parts a methods section is obliged to state plainly. Every number on this page carries its sample size and its cohort; here is what they do — and do not — mean.