Read briefing · 2026-05-13

Briefing

37 items ·
Must read

Read these first.

3 items
datacenterdynamics.com 2026-05-11 1 min read

Stream Now: DCD>Inside APAC

Brief

DCD Investment & Markets Channel's livestream “DCD>Inside APAC” (published 11 May 2026) analyzes how rapid AI investment and sovereign cloud mandates are reshaping where and how data centres are built across Singapore, Sydney, Tokyo and emerging markets in Indonesia, Malaysia and the Philippines. Key focuses include grid and land constraints, government‑industry sustainability alignment, novel deal structures to unlock capacity, and technical approaches like edge deployments and cross‑border interconnection.

By DCD Investment & Markets Channel
datacenterdynamics.com 2026-05-12 2 min read

Your streaming link: Keeping IT Cool: Liquid 2026 Day Two

Brief

Keeping IT Cool | Liquid — Day Two of DCD's Liquid series streams live at 9am ET (2pm BST) on Wednesday, 13 May 2026, focusing on liquid cooling for AI-ready data centres. Sessions cover >100 kW rack densities, scalable cooling loops, retrofit strategies, chiller-to-chip architectures, dry-break quick-disconnect reliability testing, and research from Georgia Tech, Binghamton and IBM.

By James Raddings
datacenterdynamics.com 2026-05-12 1 min read

What’s driving India’s next wave of data center expansion?

Brief

NTT Global Data Centers' whitepaper (Data Centre Dynamics, 12 May 2026) frames India as an emerging global digital‑infrastructure hub, driven by AI workloads, fast cloud adoption, data‑localization rules and growing enterprise demand. It analyzes market dynamics, infrastructure trends, investment drivers and regulatory/operational issues to help operators plan long‑term capacity, connectivity and cloud expansion.

By DCD Cloud & Hybrid
Worth reading

Deeper context and second-pass items.

12 items
divenewsletter.com 2026-05-12 1 min read

Watch Webinar | Agentic AI for Holistic Major Event Management

Brief

Agentic AI for Holistic Major Event Management (Schneider Electric, Utility Dive, 2026-05-12) presents an integrated platform approach that unifies planning, real-time operations, and post-event recovery. It emphasizes the 5–7-day pre-event window and uses AI-augmented insights, wildfire modeling, vegetation analytics, and asset-health data to produce actionable, auditable, agent-driven plans that scale across the grid lifecycle.

By Utility Dive
ArXiv 2026-05-11 1 min read

What should post-training optimize? A test-time scaling law perspective

Brief

Large-language-model post-training often optimizes single-response mean reward, creating a mismatch with best-of-N test-time selection. The paper analyzes the budget-mismatch setting (m << N) and proves that, assuming structure in reward tails, best-of-N policy gradients can be extrapolated from small-rollout samples. It introduces Tail-Extrapolated estimators (TEA, Prefix-TEA) and shows improved best-of-N performance on instruction-following benchmarks. Summary based on the abstract only.

Authors: Muheng Li, Jian Qian, Wenlong Mou
ArXiv 2026-05-11 1 min read

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge

Brief

The paper studies LLM-as-a-Judge trade-offs and finds that explicit reasoning markedly boosts accuracy on structured verification tasks (e.g., math and coding) while offering limited or negative benefit on simpler evaluations and costing substantially more compute. To address this, Wenbo Zhang et al. introduce RACER, which adaptively routes examples between reasoning and non-reasoning judges under a fixed budget via distributionally robust optimization with a KL uncertainty set. RACER uses an efficient primal–dual solver, has theoretical guarantees (unique optimal policy, linear convergence), and empirically attains better accuracy–cost trade-offs under distribution shift. (Based on the paper abstract; ICML 2026 acceptance.)

Authors: Wenbo Zhang, Lijinghua Zhang, Liner Xiang...
ArXiv 2026-05-11 1 min read

CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

Brief

CADBench is a unified multimodal benchmark for recovering editable CAD programs from 2D/3D inputs, assembled from 18,000 samples across six families and five input modalities with six evaluation metrics. The authors evaluated 11 systems (>1.4M generated programs) and found mesh-to-CAD specialists beat code-generating VLMs under ideal inputs; they report brittleness under modality shift and complexity-related failures. (Summary based on abstract; full text not available.)

Authors: Anna C. Doris, Jacob Thomas Sony, Ghadi Nehme...
ArXiv 2026-05-11 1 min read

Personal Visual Context Learning in Large Multimodal Models

Brief

Personal Visual Context Learning (Personal VCL) formalizes how wearable-driven LMMs should use user-specific first-person visual evidence to resolve personalized queries. Xue et al. introduce Personal-VCL-Bench (covers persons, objects, behaviors), identify a pronounced context-utilization gap in frontier LMMs, and propose the Agentic Context Bank—a self‑refining memory plus query‑adaptive evidence selection—that consistently improves performance over standard context prompting across tasks and evaluated backbones.

Authors: Zihui Xue, Ami Baid, Sangho Kim...
ArXiv 2026-05-11 1 min read

Factual recall in linear associative memories: sharp asymptotics and mechanistic insights

Brief

The paper analyzes limits of factual recall in linear associative memories by introducing a decoupled model and using statistical‑physics tools to obtain sharp asymptotics. It proves equivalence between the decoupled and original models (capacity, weight spectra, mechanism) and derives the capacity law pc log pc / d^2 = 1/2. The work explains why optimal learning nudges correct scores just above extreme‑value thresholds and generalizes results to two‑layer linear architectures, providing a baseline for memory in neural networks.

Authors: Alessio Giorlandino, Sebastian Goldt, Antoine Maillard
ArXiv 2026-05-11 1 min read

Count Anything at Any Granularity

Brief

Count Anything at Any Granularity (Liu et al., 2026) tackles brittle open-world object counting by making counting granularity explicit across five levels (identity, attribute, instance type, category, abstract concept). The authors build KubriCount using a novel automatic pipeline combining controllable 3D synthesis, image editing, and VLM filtering, and find existing multimodal LLMs and counting models fail to follow fine-grained prompts. They train HieraCount, which fuses text and visual exemplars, and report substantial accuracy and generalization gains on multi-grained counting tasks.

Authors: Chang Liu, Haoning Wu, Weidi Xie
Twitter/X 2026-05-12 2 min read

Associated Builders and Contractors projects the industry needs 349,000 new…

Why it matters

Associated Builders and Contractors projects the industry needs 349,000 new workers in 2026 and 456,000 in 2027; 41% of the current construction workforce will retire by 2031, while specialty-trade wages are rising 4–11% and some firms pay 20% premiums to retain crews.

Key details

  • Infill is the only practical supply path in desirable metros (SF, NYC, LA, Boston, Seattle, DC, Austin), but mid‑rise/tower infill costs 2–3× per sq ft and depends on hard‑to‑staff, union trades (ironworkers, concrete crews, complex MEP, elevator mechanics).
  • Austin permitted the most multifamily per capita from 2021–2023 and saw rents fall ~20% from the 2022 peak, but those units were built under the old code on cheap greenfield land with an ~40% immigrant workforce; Durham argues zoning alone won’t scale supply and advocates construction robotics and large prefabricated components to cut labor up to ~90%.

Brief

Nick Durham argues zoning wins won’t deliver housing quickly because the U.S. construction labor force is aging and undersized (ABC projects 349k hires in 2026, 456k in 2027; 41% retiring by 2031) and greenfield land and immigrant labor are drying up. He cites Austin’s 2021–2023 multifamily boom and 20% rent drop as driven by old‑code, low‑rise, edge‑land development, and pushes construction robotics and large prefabrication to sharply boost productivity.

By @americanhousing
ArXiv 2026-05-11 1 min read

Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with $k$-step Policy Gradients

Brief

Revisiting policy gradients for restricted policy classes, the paper identifies one-step myopia (policy gradients depending only on the one-step Q) as a cause of suboptimal critical points and proposes a k-step policy gradient that couples randomness across k steps. The method yields performance exponentially close in k to the optimal deterministic policy, converges in O(1/T) with projected/mirror descent under mild smoothness/differentiability assumptions, avoids common distribution-mismatch factors, and targets applications like state aggregation and partially observable cooperative multi-agent problems. Full text on arXiv.

Authors: Alex DeWeese, Guannan Qu
ArXiv 2026-05-11 1 min read

A Recursive Decomposition Framework for Causal Structure Learning in the Presence of Latent Variables

Brief

The paper introduces DiCoLa, a recursive decomposition framework that extends divide-and-conquer causal discovery to settings with latent variables by splitting the global CI-testing task into smaller subproblems and integrating solutions via a principled reconstruction step. The authors prove soundness and completeness for the framework and demonstrate substantial runtime improvements on synthetic experiments and practical effectiveness on a real-world dataset, addressing high-dimensional CI-testing bottlenecks.

Authors: Zheng Li, Feng Xie, Shenglan Nie...
substack.com 2026-05-12 14 min read

Introducing Barely Possible

Brief

Barely Possible is Lane Rettig’s AI-operated podcast, launched publicly on May 12, 2026 (BarelyPossible.to and on Transistor). Rettig positions the show as a high-quality, fully automated daily briefing—initially a single AI-produced program targeted at roughly one to two hours per day that synthesizes world news, AI developments, crypto and tech/finance updates, and a focused daily deep dive. He frames the product around real listening habits (he reports ~1–2 hours/day of audio and uses 2x speed to get through 3–4 episodes on a run) and wants a compact, high-octane slot that replaces time otherwise spent hunting for disparate sources.

Technically, Barely Possible runs a content-collection-and-analysis engine that synthesizes consistent underlying stories into different user-specific outputs; Rettig says the pipeline can operate on local open models, keeping costs sustainable. The project began about two months before launch with collaborator “Baz,” and Rettig has iterated on it daily. The broader vision is an AI-native audio platform with effectively infinite, customizable content (per-user “transcoding” of a single source into many tailored formats), personalized advertising, and creator tools—drawing on concepts like skill files and Neal Stephenson’s ‘ractive’—rather than attempting to supplant human-hosted long-form podcasts.

By Lane Rettig from Three Things
ArXiv 2026-05-11 1 min read

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping

Brief

Text-to-image post-training via reinforcement learning faces reward-hacking and miscalibration from normalization. Based on the abstract, the authors propose Super-Linear Advantage Shaping (SLAS), which extends the Fisher–Rao metric with advantage-dependent weighting to amplify informative updates and suppress illusory gradients, plus batch-level normalization. Evaluations report consistent gains over DanceGRPO, faster training, better GenEval and UniGenBench++ OOD performance, and improved scaling robustness and fidelity.

Authors: Haoyuan Sun, Jing Wang, Yuxin Song...
Quick skim

Fast scan items.

22 items
ArXiv 2026-05-11 1 min read

Is Your Driving World Model an All-Around Player?

Brief

WorldLens introduces a comprehensive benchmark for driving world models, spanning five aspects and 24 dimensions to evaluate pixels, 4D geometry, closed-loop driving, and human perceptual realism. The authors benchmark six models, show no method excels across axes (best attain 2–3/10 human realism), and provide WorldLens-26K plus WorldLens-Agent for scalable, explainable evaluation. (Summary based on abstract; full text not reviewed.)

Authors: Lingdong Kong, Ao Liang, Tianyi Yan...
ArXiv 2026-05-11 1 min read

Safe Aerial 3D Path Planning for Autonomous UAVs using Magnetic Potential Fields

Brief

3DMaxConvNet extends the MaxConvNet magnetic potential-field planner into 3D by training a convolutional autoencoder to generate obstacle-aware potential fields from LiDAR-derived 101^3 voxel grids. In Cosys-AirSim experiments (100 randomized closed-loop trials on two urban maps) it reached 100% success without retraining, matched A* path quality with ~2× lower runtime, and ran ~200× faster than RRT*(3k).

Authors: Haechan Mark Bong, Giovanni Beltrame
ArXiv 2026-05-11 1 min read

When Can Digital Personas Reliably Approximate Human Survey Findings?

Brief

LLM-based digital personas were evaluated as substitutes for human survey respondents using the LISS panel: personas were built from background variables and pre-2023 survey histories and tested on held-out post-cutoff answers. Across four persona architectures, three LLMs, and two prediction tasks, personas improved distributional alignment (notably for stable attributes) but struggled with individual prediction and multivariate respondent structure; retrieval augmentation helped. Summary is based on the abstract (full paper not reviewed).

Authors: Mumin Jia, Yilin Chen, Divya Sharma...
producthabits.com 2026-05-12 1 min read

Half of you got a test email yesterday

Brief

Hiten Shah (Product Habits) ran an A/B email test on 2026-05-12 — half of subscribers saw one version and half another, with some calling one variant AI-sounding. Facing a year-long drop in click rates, Shah hypothesizes readers detect “empty” AI-style prose and will run a live Reddit AMA to probe responses and defend human writing.

By Hiten Shah
Twitter/X 2026-05-12 1 min read

OpenSquilla: @hasantoxr ran a 10+ turn research workflow (published 2026-05-12)…

Brief

OpenSquilla: @hasantoxr ran a 10+ turn research workflow (published 2026-05-12) and says smart routing routed simple prompts to cheaper models and complex ones to full power, cutting his token bill by ~67% versus his usual agent while preserving memory and output quality. OpenSquilla (open-source, Python, local‑first) claims 60–80% cost reduction and launched the #10MTokenChallenge (30 winners × 10M OpenRouter credits).

By @hasantoxr
ArXiv 2026-05-11 1 min read

PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models

Brief

PriorVLA introduces a prior-preserving adaptation method for Vision-Language-Action models that freezes a Prior Expert and trains an Adaptation Expert, using Expert Queries to inject pretrained scene and motor priors. By updating only 25% of parameters, it outperforms full fine-tuning and SOTA baselines on RoboTwin 2.0, LIBERO, and eight real-world tasks, with strong OOD and few-shot gains.

Authors: Xinyu Guo, Bin Xie, Wei Chai...
ArXiv 2026-05-11 1 min read

Optimal and Scalable MAPF via Multi-Marginal Optimal Transport and Schrödinger Bridges

Brief

The paper addresses anonymous MAPF by formulating it as a Markovian MMOT that exactly reduces an exponentially large problem to a polynomial-size LP, with proven feasibility and total unimodularity yielding collision-free integral {0,1} transports. For large-scale instances the authors use a Schrödinger-bridge entropic regularization solved via Sinkhorn iterations to produce fractional templates used in a reduced LP, trading negligible optimality loss for major computational savings; experimental results demonstrate strong optimality and scalability. Accepted to ICML 2026 as a spotlight paper.

Authors: Usman A. Khan, Joseph W. Durham
ArXiv 2026-05-11 1 min read

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark

Brief

RoboMemArena is a large-scale benchmark addressing robotic memory shortcomings by offering 26 long-horizon tasks (avg >1,000 steps) with multimodal memory annotations and paired real-world evaluations; a VLM-based pipeline composes subtasks and generates trajectories via atomic functions. The authors introduce PrediMem, a dual-system VLA with recent/keyframe memory buffers and a predictive coding head that outperforms baselines on this benchmark.

Authors: Huashuo Lei, Wenxuan Song, Huarui Zhang...
ArXiv 2026-05-11 1 min read

Unified Noise Steering for Efficient Human-Guided VLA Adaptation

Brief

UniSteer tackles efficient real‑world adaptation of diffusion-based vision‑language‑action models by combining human corrective interventions with noise‑space RL. The method inverts a frozen flow‑matching decoder to convert human actions into supervision for a noise‑predicting actor, which is concurrently improved with RL. Real‑world tests on four manipulation tasks show a jump from 20% to 90% success in 66 minutes on average. Full paper was not available, summary based on the abstract.

Authors: Junjie Lu, Xinyao Qin, Yuhua Jiang...
ArXiv 2026-05-11 1 min read

Counterfactual Stress Testing for Image Classification Models

Brief

Counterfactual stress testing for medical image classifiers uses causal generative models to create realistic, anatomically-preserving ‘‘what-if’’ images by intervening on attributes like scanner type and patient sex. Evaluated on chest X‑ray and mammography across three architectures and several shift scenarios (Stammel et al., 2026), the method more reliably predicts real OOD performance than simple brightness/contrast perturbations, improving robustness assessment prior to deployment.

Authors: Moritz Stammel, Fabio De Sousa Ribeiro, Raghav Mehta...
ArXiv 2026-05-11 1 min read

Confidence-Guided Diffusion Augmentation for Enhanced Bangla Compound Character Recognition

Brief

Handwritten Bangla compound character recognition is improved via a confidence-guided diffusion augmentation approach that synthesizes class-conditional samples, enhances U-Net blocks with Squeeze-and-Excitation residuals, and filters generated images using pre-trained classifiers. Fusing filtered synthetic images with real data yields consistent gains across ResNet50, DenseNet121, VGG16 and ViT, with a top accuracy of 89.2% on AIBangla. Only the abstract was available for this summary.

Authors: Md. Sultan Al Rayhan, Maheen Islam
ArXiv 2026-05-11 1 min read

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

Brief

HarmoWAM presents an end-to-end World Action Model that unifies predictive and reactive control by conditioning a predictive expert and a reactive expert on spatio-temporal priors from a world model, with a Process-Adaptive Gating Mechanism to coordinate switching. Reported zero-shot generalization covers three unseen real-world environments and six manipulation tasks, improving over VLA and WAM baselines by 33% and 29%, respectively. Full paper text not available (abstract only).

Authors: Qiuxuan Feng, Jiale Yu, Jiaming Liu...
Twitter/X 2026-05-11 1 min read

@KayvonJafar cites Andrej Karpathy’s line “I don’t think I’ve typed a line of…

Brief

@KayvonJafar cites Andrej Karpathy’s line “I don’t think I’ve typed a line of code since December” and Garry Tan’s design prompt to have one person run a whole software team. He points to the open-source gstack as that answer, with Garry claiming ~810x higher pace; gstack wires AI roles (CEO, eng, QA, browser operator) into workflows like /office-hours, /autoplan, /qa, /review, pair-agent and parallel sprints, reframing development as operating systems of AI workers.

By @KayvonJafar
ArXiv 2026-05-11 1 min read

Price of Quality: Sufficient Conditions for Sparse Recovery using Mixed-Quality Data

Brief

Sparse recovery from mixed-quality measurements (n1 high-quality, n2 low-quality) is studied via information-theoretic and algorithmic sample-size conditions. The paper proves a linear 'Price of Quality' trade-off: agnostic decoders cap one high-quality sample at no more than two low-quality ones, whereas informed decoders can exploit variance knowledge to make the price unbounded. LASSO matches homogeneous-noise thresholds, depending only on average noise. Published ICLR 2026 (arXiv 2026-05-11).

Authors: Youssef Chaabouni, David Gamarnik
ArXiv 2026-05-11 1 min read

Pixal3D: Pixel-Aligned 3D Generation from Images

Brief

Pixal3D (Li et al., SIGGRAPH 2026) presents a pixel-aligned 3D generation paradigm that back-projects multi-scale image features into a 3D feature volume to create explicit pixel-to-3D correspondences, generating assets aligned with the input view. The method reportedly substantially raises fidelity—'approaching the fidelity level of reconstruction'—and extends to multi-view and scene synthesis.

Authors: Dong-Yang Li, Wang Zhao, Yuxin Chen...
Twitter/X 2026-05-12 1 min read

Profound's Marketing Engineering Hackathon (June 6, 2026, Union Square, NYC) will…

Brief

Profound's Marketing Engineering Hackathon (June 6, 2026, Union Square, NYC) will host 50 builders for a one-day challenge to automate inhuman-scale marketing processes, offering $40,000 total in prizes and judges from Ramp, Stripe, and MongoDB. The platform-agnostic event allows Claude Code, Cursor, n8n, Python, LangChain, or raw APIs; winners receive cash, Profound credits, and interviews. Application required.

By @hasantoxr
ArXiv 2026-05-11 1 min read

Variational Inference for Lévy Process-Driven SDEs via Neural Tilting

Brief

Variational Inference for Lévy Process-Driven SDEs via Neural Tilting introduces a neural exponential tilting framework that reweights the Lévy measure with neural nets to build a flexible variational family preserving jumps. A quadratic parametrization gives closed-form normalization; a conditional Gaussian representation for stable processes enables simulation and scalable, symmetry-aware Monte Carlo optimization. Summary based on the abstract (full text not reviewed).

Authors: Yaman Kindap, Manfred Opper, Benjamin Dupuis...
ArXiv 2026-05-11 1 min read

Geometry-aware Prototype Learning for Cross-domain Few-shot Medical Image Segmentation

Brief

GeoProto addresses cross-domain few-shot medical image segmentation by enriching prototypical matching with explicit geometric priors. The method's core, GAPE, augments appearance prototypes with learned ordinal offsets from an Ordinal Shape Branch trained under an ordinal-consistency loss (no extra labels beyond segmentation masks). According to the abstract, GeoProto attains state-of-the-art results across seven datasets and three settings (cross-modality, cross-sequence, cross-context). Only the paper abstract was available for this briefing.

Authors: Feifan Song, Yuntian Bo, Haofeng Zhang
Twitter/X 2026-05-12 1 min read

When you trade on Polymarket with PolyHelper, you automatically farm two airdrops…

Brief

When you trade on Polymarket with PolyHelper, you automatically farm two airdrops — Polymarket’s and a PolyHelper airdrop — while PolyHelper reimburses trading fees (they refunded all fees paid to Polymarket on a randomly chosen day, May 8). PolyHelper also adds extra UX info/features and is free; the author (@0xd1namit) urges there’s no reason not to use it.

By @0xd1namit
ArXiv 2026-05-11 1 min read

When Are Trade-Off Functions Testable from Finite Samples?

Brief

The paper studies finite-sample hypothesis testing for the type I/type II error trade-off function between two unknown distributions. Without structure testing is impossible; they identify exact attainability of Neyman–Pearson rejection regions by a class S and show finite VC-dimension of S is necessary and sufficient. They give a test with nonasymptotic type I control and uniform power under explicit separations, produce simultaneous confidence bands, analyze local rates (monotone likelihood-ratio) with near-matching lower bounds, and extend to approximate attainability (e.g., univariate log-concave).

Authors: Kaining Shi, Qiaosen Wang, Cong Ma
Twitter/X 2026-05-12 1 min read

CloakBrowser is a MIT‑licensed stealth Chromium (pip install, ~200MB) that the…

Brief

CloakBrowser is a MIT‑licensed stealth Chromium (pip install, ~200MB) that the author says achieves human-like anti-bot fidelity — 0.9 on reCAPTCHA v3 versus stock Playwright's 0.1 — by patching Chromium's C++ in 16 spots (canvas, WebGL, audio, fonts, hardwareConcurrency, GPU strings, WebDriver flag, TLS fingerprint). It claims to pass 14/14 detection tests, auto-resolve Turnstile, and replace commercial $200–$500+/month stealth browsers or costly custom builds.

By @hasantoxr
ArXiv 2026-05-11 1 min read

CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models

Brief

CapVector introduces a method to decouple auxiliary-objective finetuning into two parameter-space components: general capability enhancement and task-specific action fitting. By finetuning two small-scale models with different strategies and taking their parameter difference as a capability vector, then merging it into pretrained VLA models (plus an orthogonal regularizer), standard SFT attains auxiliary-like gains with lower compute. Full text not provided here.

Authors: Wenxuan Song, Han Zhao, Fuhao Li...