# Dataset manifest for MKT-004 — conjunction prediction engine for the 2026 World Cup.
# Schema: schemas/data-manifest.schema.json (EPIC §8.2).
#
# Every table is an output of the world-cup engine (github.com/away-de-petropolis/
# World-cup) at commit 63e6d53. The backtest tables are produced by the engine's
# no-leakage historical harness; the 2026 tables are produced by a single frozen
# run (snapshot c7f3f5afa805, seed 20260611). Market probabilities are devigged
# from operator-transcribed odds; their upstream provenance is noted per dataset.
paper_id: MKT-004
license: cc-by-4.0

datasets:
  - id: match-calibration
    description: |
      Held-out per-match calibration of the two team-strength engines (Elo and
      Dixon-Coles) on 1,767 international matches drawn after 2024-09-06 from an
      11,734-match window beginning in 2014. Brier, log-loss, and expected
      calibration error (ECE) on the home-win/draw/away three-way outcome.
      Match results are the martj42 international-results dataset.
    file: match_calibration.csv
    rows: 2
    schema:
      engine: enum[Elo, Dixon-Coles]
      brier: float
      log_loss: float
      ece: float
      n: int
    source: world-cup-backtest
    collection_window: [2014-01-01, 2026-06-02]
    license: cc-by-4.0
    sha256: c256ec93b1b26a122d5721c750f32599eb36275db625e2201239bb34e96cae42

  - id: champion-vs-market-history
    description: |
      Beat-the-closing-line test on the realized champion of the 2010, 2014,
      2018, and 2022 World Cups. model_p is the engine's pre-tournament champion
      probability (Elo bracket, no leakage); market_p is the devigged closing
      outright price for the same champion. Market odds are transcribed from
      vegasinsider.com outright-history tables (champion plus documented
      pre-tournament favourite only; the full free field is unavailable). The
      per-event log-losses use an overround assumption of 1.30; the sign of the
      result is robust across 1.20 to 1.40.
    file: champion_vs_market_history.csv
    rows: 4
    schema:
      year: int
      champion: string
      model_p: float
      market_p: float
      model_log_loss: float
      market_log_loss: float
      model_beats_market: bool
    source: world-cup-backtest
    collection_window: [2010-06-11, 2022-12-18]
    license: cc-by-4.0
    sha256: 224b5f3428284c6c7cc3755b58233a0256ab1acc8acc861064cb1983e5df3caf

  - id: engine-logloss
    description: |
      The engine-selection inversion. Per-match log-loss favours Dixon-Coles;
      mean champion log-loss over 2010 to 2022 favours Elo, which also beats the
      market. Dixon-Coles over-rates CONMEBOL sides because intra-confederation
      qualifiers inflate its attack ratings, so it loses on the tournament
      target despite winning per-match. The market champion log-loss (2.199) is
      the same figure as in champion-vs-market-history.
    file: engine_logloss.csv
    rows: 5
    schema:
      model: enum[Elo, Dixon-Coles, Market]
      scope: enum[per-match, champion]
      log_loss: float
    source: world-cup-backtest
    collection_window: [2010-06-11, 2022-12-18]
    license: cc-by-4.0
    sha256: 1c22303e88d378b9c7e60b1142a9a95798e646825b0fdde1b2f318969618c03b

  - id: q4-goal-count-band
    description: |
      Goal-count calibration for the Golden Boot question (Q4) across four past
      tournaments. band_5_8_mass is the model probability that the realized top
      scorer finishes on five to eight goals; realized_count_mass is the model
      probability assigned to the count that actually occurred. scorer_rank is
      the realized top scorer's rank in the model's scorer-probability list
      (blank for 2010, where four players tied on five goals). Player rates are
      the martj42 international-goalscorers dataset over a broad per-team pool.
    file: q4_goal_count_band.csv
    rows: 4
    schema:
      year: int
      top_scorer: string
      scorer_rank: int
      band_5_8_mass: float
      realized_count_mass: float
    source: world-cup-backtest
    collection_window: [2010-06-11, 2022-12-18]
    license: cc-by-4.0
    sha256: a0ea2ec4c1709c2e78ecb46e6b78a7c9f51efcb538ece1d4badadeacb3fb7442

  - id: pott-calibration
    description: |
      Player-of-the-Tournament calibration against the historical Golden Ball
      winners. model_rank and model_p are the engine's pre-tournament rank and
      probability for the realized Golden Ball winner. Calibrated coefficients
      alpha=0.5, beta=2, gamma=4 (the media vote is star-biased); mean realized
      log-loss 1.835.
    file: pott_calibration.csv
    rows: 4
    schema:
      year: int
      golden_ball: string
      model_rank: int
      model_p: float
    source: world-cup-backtest
    collection_window: [2010-06-11, 2022-12-18]
    license: cc-by-4.0
    sha256: 6a31b593f496c5f4032cc38408fc5e70d2eeb73ded3032d2889c3cdb659c5a1c

  - id: wc2026-winner-marginal
    description: |
      The twelve leading teams by the engine's 2026 champion marginal (model_p,
      Elo bracket, snapshot c7f3f5afa805) against the devigged Polymarket
      outright winner market (market_p). Probabilities do not sum to one because
      only the leading teams are listed. The model over-rates Spain and
      Argentina relative to the crowd and under-rates France, England, and
      Portugal.
    file: wc2026_winner_marginal.csv
    rows: 12
    schema:
      team: string
      model_p: float
      market_p: float
    source: world-cup-engine
    collection_window: [2026-06-11, 2026-06-11]
    fetched_at: 2026-06-02
    license: cc-by-4.0
    sha256: 2ef614fb33650bcda73076aa9dfef3b96b7034a1170c496559f6ad85f5d94e70

  - id: joint-ticket
    description: |
      The selected 2026 ticket and its nearest alternatives from a single frozen
      run (snapshot c7f3f5afa805, seed 20260611, 50,000 simulations, calibration
      gate GREEN). joint_p is the simulated probability that the entire 4-tuple
      resolves correctly. The selected ticket (goals=7) carries a joint
      probability of 0.94 percent (95 percent CI 0.86 to 1.03 percent); the
      strict argmax (goals=8) is a near-tie at 0.95 percent, and the pick is
      tilted to 7 as the less public goal count.
    file: joint_ticket.csv
    rows: 5
    schema:
      rank: int
      winner: string
      pott: string
      top_scorer: string
      exact_goals: int
      joint_p: float
      note: string
    source: world-cup-engine
    collection_window: [2026-06-11, 2026-06-11]
    fetched_at: 2026-06-03
    license: cc-by-4.0
    sha256: a52dc61536196bf35e973d67f28dd5ebc39925ec0a41d9ae2308cfcd56ab5c44

  - id: scorer-concentration
    description: |
      Concentration of the 2026 top-scorer marginal before and after integrating
      the adopted one-year recency weighting (snapshot c7f3f5afa805, 50,000
      simulations). Recency flattened the marginal: the modal share fell from
      28.5 to 23.8 percent, the effective field (inverse Herfindahl) rose from
      7.8 to 11.4 scorers, the top-five cumulative mass fell from 66.7 to 55.6
      percent, and the joint-argmax probability fell from 1.02 to 0.95 percent.
      Better current-form modelling revealed a more open Golden-Boot race.
    file: scorer_concentration.csv
    rows: 4
    schema:
      metric: string
      baseline: float
      post_recency: float
    source: world-cup-engine
    collection_window: [2026-06-11, 2026-06-11]
    fetched_at: 2026-06-03
    license: cc-by-4.0
    sha256: cf38a5d3c47fa693de68d00f7e0181d20ae668b8c94840263355ab8af1b4a17d

  - id: scorer-recency
    description: |
      Recency-weight tuning for the scorer model, mean scorer log-loss over the
      2010 to 2022 realized Golden Boots (lower is better). A one-year half-life
      improves on the all-equal baseline and was adopted. half_life_years is
      blank for the baseline arm.
    file: scorer_recency.csv
    rows: 4
    schema:
      arm: string
      half_life_years: int
      mean_scorer_log_loss: float
      adopted: bool
    source: world-cup-backtest
    collection_window: [2010-06-11, 2022-12-18]
    license: cc-by-4.0
    sha256: 541b644c6134f06a287a3f2fa89a9b5c6ed104482775fc14cb0910a93b3b5a36

  - id: scorer-recency-peryear
    description: |
      Per-tournament probability the scorer model assigned to the realized
      Golden Boot winner, all-equal baseline against the adopted one-year
      half-life. The recency arm nearly triples the probability on the realized
      winner in 2018 and nearly doubles it in 2022.
    file: scorer_recency_peryear.csv
    rows: 4
    schema:
      year: int
      golden_boot: string
      base_p: float
      recency_p: float
    source: world-cup-backtest
    derived_from: [scorer-recency]
    collection_window: [2010-06-11, 2022-12-18]
    license: cc-by-4.0
    sha256: ee83dfeeef57f968cfd98fd0871d9223543498eefa370dcce1a538714d1a1cc8
