Twitter/X

Bad Theory Labs released BTL-4 (35B) and Macaw 2.7B as open-weight models…

Brief

Bad Theory Labs released two open-weight models: BTL-4 (35B), a claimed frontier agentic reasoning model with strong metrics (73.5% BFCL v4, 78.4% SWE-bench, 66.1% LiveCodeBench) and 262K native context; and Macaw 2.7B, an on-device Mac assistant (97 macOS tools, 1.5 GB disk, ~2s per request on M2) that runs fully locally.

Why it matters

Bad Theory Labs released BTL-4 (35B) and Macaw 2.7B as open-weight models available now (Hugging Face + GitHub), pitched as a 'frontier agentic reasoning model' and an 'on-device Mac agent'.

Key details

  • BTL-4 reports 73.5% BFCL v4 (AST), +4.3 points over the base on an identical harness/decoding across 1,240 cases; 78.4% SWE-bench Verified; 66.1% LiveCodeBench v6 (99.1% easy, 86.7% medium) over a full 442-problem window; 262K native context supported, and increasing output budget 16K→32K raised LiveCodeBench by +5.2 points.
  • Macaw 2.7B claims 97 verified macOS tools chained from one sentence, 1.5 GB on disk (~2 GB running), ~2s per request on a base M2, reads your screen/documents ("uses esp as vision"), no cloud/account/telemetry, 100% local, and vLLM/MLX supported day one.
Source evidence

It didn't take two days to get the first fine-tunes of LFM2.5-2.6B. This one looks cool!

Bad theory labs (@Badtheorylabs)

Releasing BTL-4 35B and Macaw 2.7B a frontier agentic reasoning model and an on-device Mac agent. Both open weights, both available now.

• 73.5% BFCL v4 (AST) — +4.3 over the base model, paired: identical harness, identical decoding, only the weights differ, all 1240 cases

• 78.4% SWE-bench Verified
• 66.1% LiveCodeBench v6 — 99.1% easy · 86.7% medium, full 442-problem window
• 262K native context, and it uses it: raising the output budget 16K → 32K moved LiveCodeBench +5.2 points on its own. Long problems need room to finish.

Macaw 2.7B — the assistant your Mac should have shipped with. 97 verified macOS tools, chained from one sentence.
• 1.5 GB on disk, ~2 GB running, ~2s per request on a base M2 • Reads your screen and your documents — no vision model(uses esp as vision), no cloud • 100% local. No account, no telemetry, no server to send anything to

Open weights, deploy anywhere:🤗 huggingface.co/badtheorylabs… 🤗 huggingface.co/badtheorylabs… 💻 github.com/Badtheorylabs/Mac… 🔗 badtheorylabs.com/macaw

vLLM and MLX supported day one.

Link

badtheorylabs/BTL-4 · Hugging Face

We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co

— https://nitter.net/Badtheorylabs/status/2085116113881366782#m