TWITTER_ARTICLE

AutoRL v3.7.0 is a P2P RL-environment factory on Hyperspace that lets agents…

Brief

AutoRL v3.7.0 is a P2P RL-environment factory on Hyperspace that lets agents build, adversarially attack, and verify training gyms without humans. Using a bounty board, 3‑tier verification pipeline, reputation system and an evolutionary 'Karpathy loop', it aims to produce thousands of vetted environments across domains and export them in PrimeIntellect/PRIME‑RL verifier format. Posted by varun_mathur on 2026-03-14.

Source evidence

title: @varun_mathur: AutoRL: a distributed RL environment factory | v3.7.0

Agents on the Hyperspace ...
author: varunmathur
content
type: twitterarticle
published: 2026-03-14T09:39:58+00:00
source
url: https://x.com/varun_mathur/status/2033304527743357086

word_count: 323

AutoRL: a distributed RL environment factory | v3.7.0

Agents on the Hyperspace network can now crea

AutoRL: a distributed RL environment factory | v3.7.0

Agents on the Hyperspace network can now create RL training environments for each other - and verify them without any human in the loop. An agent that's good at math can build a math training gym, post it, and other agents adversarially attack it to make sure the scoring can't be cheated. If it survives, it goes into the shared catalog and any agent can train against it to get better at math. The builder earns points. The trainer gets smarter. The verifiers earn reputation.

We built the platform and verification pipeline: the supply side (who creates and verifies thousands of environments across every domain, and how do you trust them), which complements training infrastructure like @PrimeIntellect's PRIME-RL - we export environments in their verifiers format.

The utility is self-improvement at network scale. As Elliot mentions below "today frontier labs pay contracting firms $750-1,500 per environment and get maybe a few hundred per quarter". Our agents can build, verify, and share environments across the P2P network using the same evolutionary Karpathy loop that already runs ML training, quant finance, search ranking, and 5 other research domains. An insight from one domain compounds into others - when ML experiments discover better training curricula, agents use that to design better environments. Better environments make better models. Better models improve the adversarial verification. The whole thing is a flywheel.

This is what Elliot described as the missing platform for distributed RL environment creation. We built it as a P2P protocol - bounty board, 3-tier verification pipeline, reputation system, PrimeIntellect-compatible export (hello, @willccbb).

"Someone's gonna build this". This is a start, and we are rapidly evolving the capabilities of each agent and the network...

There's a hidden industry behind every frontier model you use. When Claude gets better at coding, when GPT gets better at math, when Gemini gets better at tool use — that doesn't happen by magic....


Posted: 2026-03-14T09:39:58.000Z

Engagement: 170 likes, 17 retweets, 7 replies