Twitter/X

Levie (2026-08-02) argues that objectively verifiable, high‑complexity…

Brief

Levie (2026-08-02) argues that some of the hardest, highest‑value work—specifically math, cyber, and code—will be automated earlier than less verifiable work because those domains yield objective tests and clear training rewards, enabling scalable validation of model outputs. He contrasts that with legal negotiation, marketing strategy, sales messaging, and financial planning, which have no single right answer, depend on operator risk/preferences and context, and often only show results after significant delay. The implication is that automation gains will come more from applied‑AI integration and process redesign than from models alone, and that new methods to test knowledge work over time will be needed. Max Spero adds a three‑tier verifiability framework: programmatic (fastest), real‑world/time‑bounded (sim‑to‑real limits progress), and human‑preference (culturally shifting, unlikely to reach stable superhuman performance).

Why it matters

Levie (2026-08-02) argues that objectively verifiable, high‑complexity fields—math, cybersecurity, and coding—are prone to earlier automation because testable outputs provide clearer reward signals for model training and scalable verification of correctness.

Key details

  • Levie identifies domains resistant to early automation—legal clause negotiation, marketing campaign selection, sales messaging, and business financial planning—because they lack a single right answer, depend on operator risk/preferences and context, and often only reveal correctness after long delays.
  • Levie concludes that capturing automation gains will require substantial work at the applied‑AI layer and process redesign, plus new capabilities to “test” knowledge work over time similar to software testing.
  • Max Spero categorizes verifiability into three tiers: (1) programmatically verifiable (games, coding, math, cybersecurity, chip design) likely solved fastest; (2) real‑world/time‑bounded (biology, chemistry, materials, pharma, aerospace, robotics, forecasting) limited by sim‑to‑real and time/cost; (3) human‑preference (writing, design, comedy, persuasion) probably never fully solved because cultural preferences keep shifting.
Source evidence

We’re going to be in for a strange dynamic which is that some of the “hardest” work in the world is actually prone to automation first, particularly due to its verifiability.

Math, cyber, and code -while being insanely hard and high value fields- have the benefit of being able to be tested that it’s correct objectively. This has two immediate benefits: the training of the models offers clearer reward signals, and then the running of the models allows you to know that it’s working properly because you can test the results in a scalable way.

Conversely, in other domains of work, there’s much less instant verifiability. Which legal clauses your client will agree to, what marketing campaign to run with based on changing sentiment, which message your sales prospect will want to hear, what financial targets and budget to set for a business, and so on.

All of these domains have changing internal and external factors, they don’t have “one right answer”, they rely on the opinions and risk levels of the operators, they’re highly sensitive to getting the right input context first, and in many cases the right answer can’t even be known for quite some time after the model generates the results.

The implications of this distinction are that -even as model capability continues to increase exponentially- there will be a lot done at the applied AI layer than just the the model itself, and much of the processes themselves will even need to change over time to get the full gains from automation. We may even need all new capabilities to be able to “test” knowledge work over time as we have had with software.

Max Spero (@maxspero)

In my view we have a few different tiers of verifiability
1) programatically verifiable (near-free)
- games, coding, math, cybersecurity, chip design
2) real-world verifiable (cost-or time bounded)
- sciences: biology, chemistry, physics
-physical world: material science, energy, aerospace, robotics, agriculture, pharma
-forecasting: trading, weather
3) verifiable with human preference
- writing, design, comedy, charisma, persuasion

I expect most low-hanging fruit in (1) to be solved very quickly. Not sure how much longer before more solved math conjectures are simply uninteresting. I’m least certain about chip design being in this category, perhaps we hit a ceiling and require physics or materials breakthroughs to continue progress.

My guess is we will quickly run into the limits of how well we can simulate each domain in (2). The time-bounded nature of real world verification may be the reason we don’t hit fast takeoff. Sim2real remains an elusive problem to solve when real-world data is limited. Part of the reason I don’t expect to live multiple hundreds of years is simply that I expect pharmaceutical progress to be time-bounded by the physical world.

I expect the items in (3) to never really get solved to a superhuman degree, as success relies on an ever-shifting plane of cultural preference. People adapted to “good” AI writing and became annoyed at new stylistic tics that, in a vacuum, are not necessarily bad.

— https://nitter.net/maxspero/status/2083600597253521745#m