Model cards · engine v14One word per market.
One word per market.
Computed, never written.
Every market carries one of five statuses. It is computed from the committed evidence against rules fixed before the results were read, and the build fails if anyone edits a label by hand. Cards generated Sep 25, 2026.
Tier 1PromisingPitcher strikeoutsIt passes the backtest. The live forward test is still running: 0 of 200 settled results over 0 of 7 dates.25.8%lower error than the league rate · 22.5% to 28.6%Tier 2Promising1+ hitIt passes the backtest. The live forward test is still running: 0 of 200 settled results over 0 of 7 dates.1.4%lower error than the league rate · 0.9% to 1.8%Tier 2Promising2+ total basesIt passes the backtest. The live forward test is still running: 0 of 200 settled results over 0 of 7 dates.1.2%lower error than the league rate · 0.9% to 1.6%Tier 3Research only1+ home runResearch only by policy. It passes the backtest. The live forward test is still running: 0 of 200 settled results over 0 of 7 dates.1.5%lower error than the league rate · 1.1% to 1.9%Tier 3Research onlyGame totalResearch only by policy. In the backtest it is not measurably more accurate than the league-wide rate.-0.7%lower error than the league rate · -3.2% to 1.8%Tier 3Research onlyMoneylineResearch only by policy. In the backtest it is not measurably more accurate than the league-wide rate.1.3%lower error than the league rate · -0.4% to 2.9%Tier 3Research onlyFirst inning (NRFI/YRFI)Research only by policy. The backtest has 0 of 500 graded predictions and 0 of 60 dates.No backtest yet
The five words
| Validated | Beats the league-wide rate in the backtest, is not worse than the player's own record, is calibrated, and live results under the engine now serving confirm it. |
| Promising | Real skill in the backtest, but one check is still open: calibration spread, or the live forward test. |
| Insufficient sample | Not enough graded history to judge it. |
| Underperforming baseline | A constant or the player's own track record does as well or better. |
| Research only | Tier 3 by policy. Published for context, never as a basis for a bet, until a deliberate product change promotes it. |
The rules
Fixed on September 17, 2026, before the results were read.
- Backtest
- At least 500 graded predictions over 60 dates, replayed walk-forward.
- Skill
- The 95% interval of skill over the league-wide rate is entirely above zero.
- Player baseline
- Not significantly worse than the player's own shrunk record.
- Calibration
- Calibration slope of the served number between 0.8 and 1.2.
- Live test
- At least 200 settled pre-game results over 7 dates under the engine now serving, with positive skill.
The comparison with the book's price is on every card and changes no status. It answers a different question: is this number a reason to bet?
Tiers
- Tier 1
- Lead market. Promoted to Validated when the evidence says so.
- Tier 2
- Shown with caution.
- Tier 3
- Research only until proven.