Home Betting Model Methodology Track Record Prop Betting Pitcher Evaluation Barrel Rate Guide 2026 Projections Season Preview Division Predictions Free Agency Offseason Hub Analytics Hub xFIP Guide wOBA Guide BABIP Guide Daily Analysis Beyond Box Score MLB Trends About 🎯 Free Picks Today

The Start Context Model: Pricing Three Games From What Actually Scored

August 27, 2026 · Share on X · RSS · Email

Almost every public run model starts from a pitcher's ERA and a team's runs per game, then applies a park factor and calls it a projection. That pipeline has a structural flaw. ERA measures what the starter allowed, and a baseball bet settles on what the whole team allowed. The two numbers differ by ten to twenty percentage points at the tails, and every projection built on the first one inherits that bias.

The Start Context Model replaces the first input entirely. Instead of asking how a starter performed, it asks a blunter question with a directly observable answer: when this pitcher started, what did the opposing club actually finish with on the scoreboard? Runs, not earned runs. Nine innings, not five and two thirds. That single figure already contains the bullpen, the defense, the manager's hook and every extra inning the starter did not pitch.

The Two Inputs

The model has exactly two inputs and one scaling term. It is deliberately small, because a small model that can be audited beats a large one that cannot.

Input one, start context. For each starting pitcher, the mean runs scored by the opposing club across the full games he started in 2026. Pulled from his game log, resolved game by game to the final line score.

Input two, team run prevention. The starter's club runs allowed per game across the full season. This is the regression term. A starter's own sample is 13 to 24 games and will overstate both the good and the bad, so the model blends the two halves evenly rather than trusting either alone.

The scaling term. The opposing offense, expressed as its runs per game divided by the 2026 league average of 4.467. A club that scores 4.96 gets a 1.11 multiplier. A club that scores 4.14 gets 0.93.

Expected runs = ( 0.5 × start context + 0.5 × team runs allowed per game ) × ( opponent runs per game / 4.467 )

Convert the two expected run figures into Poisson distributions, convolve them, split the ties evenly, and the output is a win probability and a full scoring distribution for each club. From that distribution any team total threshold can be priced directly.

Tonight's Inputs

Six starters, three games. The gap between the raw start context column and the team column is where the regression does its work, and it is largest exactly where you would expect it: on the two pitchers with the smallest samples.

StarterStartsStart contextTeam runs allowedBlended
Jacob Misiorowski242.633.713.17
Sean Manaea134.694.484.59
Gerrit Cole162.813.723.27
Hayden Wesneski53.804.934.37
Yoshinobu Yamamoto232.573.853.21
Chris Sale232.393.853.12

Notice what the regression does to Manaea. His raw start context of 4.69 is the worst on the board, but the Mets allow 4.48 as a club, so the blend barely moves. That is the model saying his number is not a small sample artifact. It is who New York is on the days he pitches.

Truist Park in Atlanta, host of the Dodgers and Braves game the model prices closest to a coin flip

Truist Park. The model separates the Dodgers and Braves by three tenths of a run and calls the game 51.5 to 48.5. Photo: Andrew nyr, Wikimedia Commons, CC BY-SA 4.0

The Output

Run the three games through the equation and the model produces an expected run figure for each club, a win probability, and a fair moneyline. The fair line is the price at which the model would be indifferent, before any vig.

GameExpected runsModel probabilityModel fair lineMarket
Brewers at Mets5.09 to 2.94Milwaukee 77.5%-344-191
Astros at Yankees4.41 to 3.34New York 64.7%-183-151
Dodgers at Braves3.47 to 3.37Los Angeles 51.5%-106-129

Two of the three disagree with the market in the same direction. The model wants Milwaukee at -344 and can buy it at -191, and it wants New York at -183 and can buy it at -151. Those are large gaps and large gaps deserve suspicion rather than celebration, which the limitations section takes up below.

The Atlanta Team Total

Because the model outputs a full Poisson distribution rather than a point estimate, any threshold on either club can be priced without extra work. Atlanta's expected runs against Yamamoto and the Los Angeles staff come out at 3.37.

Atlanta outcomeModel probabilityFair price
2 runs or fewer34.6%+189
3 runs or fewer56.5%-130

That is the most interesting output of the night, because it lands between the two prices available on the same bet. The Atlanta under 3.5 is posted at -145 on the card BetLegend published, and the TrustMyRecord market board carries the same selection at -120. The model says the true number is about -130. At -120 the under is a position. At -145 it is a lean the model would not take on its own.

None of the three lines above is a recommendation. Every play mentioned on this site comes from BetLegend's posted card and is graded publicly at TrustMyRecord. The model exists to price those plays honestly, including when it disagrees with them.

Where This Model Is Weak

A model presented without its failure modes is marketing. These are the four that matter most tonight.

Poisson understates baseball's tails. Real run distributions are overdispersed, because innings are not independent. A Poisson fit will systematically underprice both shutouts and ten run games. For a threshold sitting near the mean, like Atlanta at 3.5 against an expectation of 3.37, that error is small. For a threshold far out in the tail it would not be.

Wesneski has five starts. The regression pulls his 3.80 most of the way toward Houston's 4.93 for exactly that reason, but five starts is not a sample, and the Yankees number carries the widest error bar of the three.

No park factors, no weather, no lineups. The model uses season aggregates. It does not know that Citi Field suppresses right handed power or that a lineup card had not posted when this ran. Those are real effects and their absence is a known bias, not a rounding error.

Start context is partly the offense's fault. Runs allowed in a pitcher's starts include the innings where his own club was up seven and the leverage arms sat. Blowouts inflate the figure for good teams. This is the single largest unmodeled distortion in the input.

Rebuilding It Yourself

Every number above comes from two public endpoints. Pull the 2026 game log for a starter, filter to games he started, and take the gamePk for each. Resolve each gamePk to a final line score and take the runs for the club he faced. Average those. That is start context. Divide the club's season runs allowed by games played for the second input, and take the opposing offense's runs per game divided by 4.467 for the multiplier. The rest is a Poisson convolution in twenty lines.

The reason to build it this way is not accuracy for its own sake. It is that when the output disagrees with a price, you can point at the exact input responsible. A model you cannot argue with is a model you cannot use.

More from Prediction Models

Related Analysis

Starting Pitcher Evaluation Pitching Analytics 2026 xFIP Explained for Betting BABIP Regression Guide wOBA Betting Applications