EdgeK.ai
About
EdgeK

Calibrated. Tracked. Transparent.

A free MLB pitcher strikeout prediction model that shows its work — every prediction, every result, every miss.

What makes EdgeK different

What EdgeK does

Every day, EdgeK pulls the day's probable starting pitchers straight from the MLB schedule and scores every one of them — predicted strikeout count, an 80% confidence interval, and a talent tier — on the Predictionstab. Nobody gets left out for any commercial reason; it's every starter, every day.

Predictions settle automatically against real box scores the next morning. Yesterday's slate, with actual strikeout counts and per-game error, lives on Yesterday. The full historical accuracy — MAE, bias, hit-rate bands, calibrated confidence intervals — is on Accuracy, and it grows by one day's slate every single day the model runs.

How the model works

Strikeout count = K-rate per plate appearance × batters faced. EdgeK models both halves separately: a gradient-boosted Stage A predicts a pitcher's point-in-time K% from Statcast pitch-level data (whiff rate, chase rate, pitch mix, velocity), and a second model, Stage B, predicts how many batters they'll actually face — the workload half most K-count models skip entirely.

Strikeout counts per start follow a Negative Binomial distribution — same family as Poisson, but with the spread tunable per pitcher. Each pitcher gets their own dispersion estimate (α), empirical-Bayes-shrunk toward the league prior when their own sample is thin, so a rookie with three starts doesn't get an overconfident interval.

The raw probabilities get a final Beta calibration pass to keep them honest, and the prediction intervals get a conformal recalibration against the full backtest — so when the model says "80% range," empirical coverage actually hits close to 80%, not the 88%+ a naive interval would over-cover at.

Why no more betting

EdgeK ran as a live value-betting model for several months. The honest result, measured the same rigorous way everything else on this site is measured: the model's point predictions were genuinely good — well-calibrated, tracking near the statistical noise floor for K-count prediction — but the betting edge (how often the model's disagreement with the market actually paid off) collapsed toward zero. Sportsbooks price strikeout props about as sharply as any public model can beat, and that stopped being a productive place to spend effort.

Rather than keep publishing a betting product built on a edge that had evaporated, EdgeK dropped the entire odds/market layer — no more lines, no more edge threshold, no more ROI. What's left is the part that was always genuinely good: a well-calibrated prediction model, published as a free, transparent tool rather than a bet recommendation engine.

The numbers right now

Predictions
3,554
self-updating daily
MAE
1.96
Ks off, on average
Bias
+0.24
predicted − actual
Within ±1K
46%
of predictions
Full hit-rate bands and calibrated CI coverage on the Accuracy tab.

Methodology choices

  • Full-slate coverage. Every probable starter gets scored every day, sourced straight from the MLB schedule — no market-line gating. Earlier versions of this pipeline silently skipped pitchers without a betting line; that gap is gone.
  • Void on no official start. A prediction is excluded from accuracy stats if the pitcher didn't get the official start that day — including a bulk reliever coming in behind an opener.
  • Per-pitcher variance. Dispersion (α) is fit per pitcher from their own start-to-start history, empirical-Bayes-shrunk toward the league prior for thin samples, so confidence intervals reflect each pitcher's real volatility, not a one-size-fits-all band.
  • Talent tiers, half-life weighted. Each pitcher's tier (scrub/avg/good/elite) is a 1-year half-life-weighted average of career strikeout rate — recent form counts more than a stale season-old number.

What this is (and isn't)

EdgeK reports model output for research and informational purposes. It is NOT betting advice, financial advice, or a guarantee of future accuracy. Past performance — even honestly calibrated past performance — doesn't promise tomorrow's results.