An Introduction to MLB Handicapping Models for UK Bettors

Updated July 2026
Licensed
Available in US
Fast payouts
18+ Only
An Introduction to MLB Handicapping Models for UK Bettors
Last updated: Reading time : 11 min

The first MLB handicapping model I built ran on a battered laptop and a spreadsheet with about 40,000 cells. It produced projections for the day’s slate that were, in retrospect, only marginally better than reading the games off the back of a newspaper. What it did teach me was that the act of building a model – of having to specify exactly what inputs you believe matter and how much – forces a clarity of thought that no amount of staring at the lines produces.

Most UK punters who progress past casual betting on MLB eventually face the question of whether to build or borrow a handicapping model. The answer is more nuanced than the keen-amateur enthusiasm or the dismissive scepticism would suggest. This guide walks through what handicapping models actually do, what they cannot do, and how a disciplined UK punter should think about integrating modelled outputs into their workflow.

What a Handicapping Model Actually Produces

The output of a good MLB handicapping model is not a pick. It is a probability estimate. The model takes the inputs available – lineups, starting pitchers, bullpen state, park factor, weather, umpire, recent form – and produces a number that estimates the home team’s probability of winning the game. That number, combined with the line at the operator, determines whether a wager carries positive or negative expected value.

A model that says the home team has a 55 percent win probability, against a market price of 1.85 in decimal (implying 54.05 percent), gives a small positive edge on the home moneyline. The same model output against a market price of 1.75 (implying 57.14 percent) gives a small negative edge – the model thinks the home team’s probability is lower than the market implies, so the home moneyline is overpriced.

This framing matters because most casual punters think of betting as picking winners. A model thinks of betting as identifying mispriced probabilities. Across a season of meaningful volume, picking winners is essentially impossible to do at a rate that consistently beats the books’ hold. Identifying mispriced probabilities – even at modest accuracy – is the only known path to profitable wagering.

The Inputs That Actually Matter

A working MLB handicapping model has somewhere between 8 and 20 core inputs. Below that count, you’re missing variables that meaningfully affect outcomes. Above that count, you’re adding inputs that produce no additional predictive signal but do introduce overfitting risk.

The essential inputs are: starting pitcher quality and form, lineup composition with platoon adjustments, bullpen quality and current availability, park factor for the venue, weather conditions during the game window, umpire’s strike zone tendencies, recent run-scoring form for both teams, and home-field advantage adjustment. Each of these has been covered in dedicated articles across the cluster. The article on pitcher versus hitter matchups goes into the matchup-input mechanics in more detail.

The inputs that most casual punters over-weight: career batter-versus-pitcher histories (sample size too small), recent winning or losing streaks at the team level (mostly noise around true talent), media narratives about momentum (no replicable signal). The inputs that get under-weighted: bullpen state, umpire effects, weather effects, and lineup-context adjustments. A model that gives appropriate weight to the second list and dismisses the first list is already operating ahead of most casual betting analysis.

How Models Compare to the Closing Line

The benchmark for any handicapping model’s quality is whether it produces probability estimates that beat the closing line at the major operators. The closing line is the market’s final integrated estimate of the game’s probability after all available information has been priced in. A model that consistently produces estimates closer to the eventual game outcomes than the closing line implies has genuine predictive power.

The honest reality is that very few handicapping models meaningfully beat the closing line across long samples. The closing line is the consensus of all professional money, including dedicated MLB-only handicappers running far more sophisticated models than any individual UK punter is likely to build. The structural edge versus the closing line is small.

What models can do – and what makes them worth building – is beat the opening line and the mid-cycle prices at smaller operators. The closing line is the consensus, but the prices through the line’s lifecycle reflect varying levels of market efficiency. A model that beats the opening line by 1 to 2 percent on probability estimates can produce a profitable strategy if the punter places wagers early when the model identifies edges versus opening prices.

The Spreadsheet Approach for UK Punters

For UK punters who want to start with modelling but who are not professional programmers, a working spreadsheet-based model is achievable in a weekend of focused work. The structure is straightforward – one sheet for player statistics, one sheet for team statistics, one sheet for game-level inputs (weather, umpire, park), one sheet that combines everything into a probability output.

The data sources are publicly available. MLB statistics are mirrored across multiple sites with free APIs. Weather data is accessible through public forecasting services. Umpire assignments are announced morning-of and aggregated by independent sites. Park factors are stable enough to be hardcoded once and updated rarely.

The complexity is in the formulas that combine these inputs into a probability estimate. The simplest working approach is a linear combination – each input contributes a weighted adjustment to a baseline probability, and the weights are calibrated against historical results. More sophisticated approaches use logistic regression or machine learning, but those require more programming skill and are not always materially better than a well-calibrated linear model.

What Models Cannot Do

The limits of modelled outputs are as important to understand as their strengths. A model cannot predict individual game outcomes with high confidence – the variance in baseball is too high. A model that’s correct on probability estimates can still produce a losing streak of 10 or 15 games purely from variance. The structural reality that around 30 percent of MLB games are decided by one run means that even the best-priced wagers carry enormous swing-game exposure, and the punter who has not internalised that mathematics will abandon a working model during a normal bad run.

A model cannot account for information it does not have. A clubhouse rumour, a pre-game injury that has not been reported, a manager’s tactical decision that breaks pattern – these are inputs the model does not see and therefore does not price. The closing line integrates them through the price action of bettors who do see them. The model often does not.

A model cannot tell you when to stop trusting it. A model that’s working well in April might be miscalibrated by July as the league’s run environment shifts or as your input weights drift out of correspondence with reality. The disciplined practice is to monitor the model’s calibration continuously and update it when the predicted versus actual outcomes diverge consistently.

The 2025 Market Context for Modelling

The structural environment for MLB handicapping has changed meaningfully across the past several seasons. The 53 percent of US bettors who wagered on MLB in 2025 are operating in a market with deeper liquidity than ever, more efficient closing lines on the major US operators, and tighter integration of public data into the books’ own modelling.

For UK punters, this means the structural edges versus the closing line at UK-facing books are smaller than they were a decade ago, but the opening-line and mid-cycle edges remain meaningful. The hold percentage on US sports betting markets averages 10.15 percent in 2025 overall, with MLB-specific holds in the 4 to 6 percent range – which is the structural barrier that any modelling approach has to clear to be profitable.

The good news is that the public-data infrastructure for MLB is the most comprehensive of any major sport. Every pitch is tracked, every batted ball is logged with exit velocity and launch angle, every defensive positioning is recorded. The raw materials for a working model are abundant. The discipline is in selecting which inputs to use and how to combine them, not in finding the inputs.

When to Build, When to Borrow

The build-versus-borrow question depends on the punter’s resources, technical skills, and time horizon. Building a working model is roughly 30 to 60 hours of focused initial work plus 1 to 2 hours per week of maintenance through a season. Borrowing – using publicly available projections from established systems and combining their outputs with your own judgment – is roughly 0 hours of initial work and 30 minutes per day of integration.

For most UK punters who are progressing past casual betting, borrowing is the rational starting point. The publicly available MLB projection systems (FanGraphs, Baseball Prospectus, and others) produce daily output that is competitive with most amateur-built models, often better. Using their projections as a baseline and overlaying your own adjustments for current information is a workable approach that takes a fraction of the time of building from scratch.

Building becomes worthwhile when you have specific insights you believe are not reflected in the public projections. A specialised model for park factor adjustments, for example, or for late-game bullpen states, might produce genuine edge that combines productively with the publicly available baseline projections. Building a full system from scratch is rarely the optimal use of a UK punter’s time.

The Integration Discipline

Whatever the source – built model or borrowed projections – the discipline that distinguishes profitable from unprofitable punters is the integration between model outputs and actual wagering decisions. The model is one input among several. The punter’s judgement on game-specific factors that the model does not capture is another. The market line is a third.

The right approach is to use the model output as a probability estimate, compare it to the market’s implied probability, and only wager when the gap is large enough to overcome the hold and the punter’s confidence in the estimate. A 1 percent gap is rarely enough to wager on. A 3 percent gap is usually enough. The size of the wager scales with the size of the gap, capped at a fraction of the bankroll determined by overall risk tolerance.

Modelling and Workflow Questions UK Punters Ask

The two questions below cover the cases that come up most often when newer modellers start integrating outputs into their betting workflow.

Can a hobbyist MLB model genuinely produce profitable wagering?

It can, but the edges are small and the discipline requirements are substantial. Most hobbyist models produce marginal edges versus opening lines at smaller operators. Reaching consistent profitability requires both a working model and disciplined bankroll management.

Should I trust publicly available MLB projections or build my own?

Start with publicly available projections. They are typically competitive with hobbyist-built models. Build your own when you have specific insights that the public projections do not capture and that you can validate against historical data.

This material was created by the Mound & Margin team.

Related posts