Model cards · engine v14

One word per market.
Computed, never written.

Every market carries one of five statuses. It is computed from the committed evidence against rules fixed before the results were read, and the build fails if anyone edits a label by hand. Cards generated Sep 25, 2026.

The five words

ValidatedBeats the league-wide rate in the backtest, is not worse than the player's own record, is calibrated, and live results under the engine now serving confirm it.
PromisingReal skill in the backtest, but one check is still open: calibration spread, or the live forward test.
Insufficient sampleNot enough graded history to judge it.
Underperforming baselineA constant or the player's own track record does as well or better.
Research onlyTier 3 by policy. Published for context, never as a basis for a bet, until a deliberate product change promotes it.

The rules

Fixed on September 17, 2026, before the results were read.

Backtest
At least 500 graded predictions over 60 dates, replayed walk-forward.
Skill
The 95% interval of skill over the league-wide rate is entirely above zero.
Player baseline
Not significantly worse than the player's own shrunk record.
Calibration
Calibration slope of the served number between 0.8 and 1.2.
Live test
At least 200 settled pre-game results over 7 dates under the engine now serving, with positive skill.

The comparison with the book's price is on every card and changes no status. It answers a different question: is this number a reason to bet?

Tiers

Tier 1
Lead market. Promoted to Validated when the evidence says so.
Tier 2
Shown with caution.
Tier 3
Research only until proven.