Odd Lots

Why Soccer Analytics Works Like Volatility Arbitrage Trading

Brief

Soccer analytics, the guests argued, has moved from the "too chaotic for moneyball" camp into a data- and compute-driven field that looks and feels a lot like volatility arbitrage in finance. The episode used the 2026 World Cup as a lens: hosts cite Morgan Lewis numbers that 104 matches will produce over 90 petabytes of match-based data (a ~45x rise from 2022), and both guests — Mike (Apex Fintech head of risk, ex-volatility trader) and Yours (professional soccer analytics consultant) — mapped how that data explosion enables deeper models. Mike framed soccer outcomes as distributions with high variance (promotion/relegation compounds financial volatility), so clubs must manage player portfolios rather than simply buy the best names. He described MLS roster mechanics as a clear example: slot-based salary-cap charges mean a designated player's cap hit can be much smaller than his salary, forcing relative-value decisions across the roster.

Technically, the conversation traced the arc from on‑ball event data (row-by-row events with X/Y/time) through positional tracking at ~10–25 frames/sec to full skeletal/body-pose data with ~27 keypoints per player — the source of the petabyte-scale volumes. Yours described how model outputs (classification windows, expected possession value/expected threat) are rarely fed raw to coaches: analysts must translate signals into short video clips that are actionable in pre-match or half-time windows. Both guests acknowledged limits: VAR, referee errors and rare game states are treated as uncontrollable noise and often censored; analysts must do significant data hygiene to avoid overfitting (Mike gave the Nicolas Jackson example where an anomalous 3‑goal match skewed seasonal stats). They also debated distributional effects: open-source tools can help smaller clubs, but access to compute and talent could concentrate advantage. The hosts ended on broader implications — parallels to chess and Goodhart-like pitfalls — stressing that analytics asks "what are you optimizing for?" (wins, profit, avoiding relegation), and that measurable gains depend on careful modeling, translation, and organizational buy-in rather than raw data alone.

Why it matters

Morgan Lewis figures (cited by Joe Wisenthal) project the 2026 FIFA World Cup's 104 matches will generate >90 petabytes of match-based data — a ~45x increase over the 2022 World Cup.

Key details

  • Mike (head of risk at Apex Fintech, ex-volatility arbitrage trader) argues soccer is fundamentally a set of distributions: team and player performance are high-variance, amplified by promotion/relegation, which creates a portfolio-management problem for clubs.
  • Mike explains MLS roster rules force portfolio thinking: each roster slot carries a salary-cap charge (e.g., a designated player's cap charge can be far lower than his salary — Mike cited an example where a superstar's cap charge could be ~£750k while actual pay is much higher), so clubs must allocate finite cap space across player slots.
  • Yours (soccer analytics consultant) and Mike described the data stack: legacy 'on-ball' event data are row-based events with X/Y/time tags; modern tracking data record all players at ~10–25 frames/sec and skeletal/body-pose data add ~27 body keypoints per player (orders of magnitude more data).
  • Both guests emphasized model limitations and data hygiene: irregular game states (red cards, teams reduced to nine) are often censored because they skew season aggregates — Mike used a Tottenham match where Nicolas Jackson scored 3 goals (≈20% of his season total) as an example.
  • Translation and workflow: Yours stressed teams need analysts to convert opaque ML outputs into coach-actionable items (typically video clips). Most analytics-driven tactical changes occur pre-match, at half-time, or during official breaks — truly real-time autonomous substitutions based solely on live ML signals are rare.
Reader · no content

No body text on file.

Open the original to read the full piece.