Twitter/X

Rui Wang (@RRicefan) tested state-of-the-art LLMs over a two-year simulated…

Brief

Rui Wang (@RRicefan) reports a two-year evaluation of SOTA LLMs on stock trading: models underperformed a simple static baseline, increased reasoning (e.g., chain-of-thought) did not help, and at least one agent (Sol) responded to losses by trading less instead of correcting strategy, showing weak practical trading behavior.

Why it matters

Rui Wang (@RRicefan) tested state-of-the-art LLMs over a two-year simulated trading period and concludes they "can't trade" and "suck."

Key details

  • No model came close to a simple static baseline; adding more chain-of-thought or higher reasoning did not improve trading performance.
  • When losses occurred, the Sol agent reduced trading activity rather than adapting to improve results, indicating poor loss-recovery behavior.
Source evidence

This is a pretty deep analysis of what happens when you try to get LLMs to trade stocks.

Rui Wang (@RRicefan)

LLMs can't trade & higher reasoning doesn't help.

we ran SOTA models for a 2y period. TL;DR: they suck & reasoning doesn't help.
- no model comes close to simple static baseline
- more reasoning ≠ better trading
- when losing money Sol trades less instead of better

details 👇

— https://nitter.net/RRicefan/status/2082513323489202664#m