Briefing · 2026-05-13

Your briefing

37 ranked ·

Today's dispatch

Filed · 37 ranked

  1. 95 score datacenterdynamics.com · Must read · 1 min Stream Now: DCD>Inside APAC DCD Investment & Markets Channel's livestream “DCD>Inside APAC” (published 11 May 2026) analyzes how rapid AI investment and sovereign cloud mandates are reshaping where and how data centres are built across Singapore, Sydney, Tokyo and emerging markets in Indonesia, Malaysia and the Philippines. Key focuses include grid and land constraints, government‑industry sustainability alignment, novel deal structures to unlock capacity, and technical approaches like edge deployments and cross‑border interconnection.
  2. 93 score datacenterdynamics.com · Must read · 2 min Your streaming link: Keeping IT Cool: Liquid 2026 Day Two Keeping IT Cool | Liquid — Day Two of DCD's Liquid series streams live at 9am ET (2pm BST) on Wednesday, 13 May 2026, focusing on liquid cooling for AI-ready data centres. Sessions cover >100 kW rack densities, scalable cooling loops, retrofit strategies, chiller-to-chip architectures, dry-break quick-disconnect reliability testing, and research from Georgia Tech, Binghamton and IBM.
  3. 92 score datacenterdynamics.com · Must read · 1 min What’s driving India’s next wave of data center expansion? NTT Global Data Centers' whitepaper (Data Centre Dynamics, 12 May 2026) frames India as an emerging global digital‑infrastructure hub, driven by AI workloads, fast cloud adoption, data‑localization rules and growing enterprise demand. It analyzes market dynamics, infrastructure trends, investment drivers and regulatory/operational issues to help operators plan long‑term capacity, connectivity and cloud expansion.
  4. 82 score divenewsletter.com · Worth reading · 1 min Watch Webinar | Agentic AI for Holistic Major Event Management Agentic AI for Holistic Major Event Management (Schneider Electric, Utility Dive, 2026-05-12) presents an integrated platform approach that unifies planning, real-time operations, and post-event recovery. It emphasizes the 5–7-day pre-event window and uses AI-augmented insights, wildfire modeling, vegetation analytics, and asset-health data to…
  5. 82 score ArXiv · Worth reading · 1 min What should post-training optimize? A test-time scaling law perspective Large-language-model post-training often optimizes single-response mean reward, creating a mismatch with best-of-N test-time selection. The paper analyzes the budget-mismatch setting (m << N) and proves that, assuming structure in reward tails, best-of-N policy gradients can be extrapolated from small-rollout samples. It introduces…
  6. 79 score ArXiv · Worth reading · 1 min Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge The paper studies LLM-as-a-Judge trade-offs and finds that explicit reasoning markedly boosts accuracy on structured verification tasks (e.g., math and coding) while offering limited or negative benefit on simpler evaluations and costing substantially more compute. To address this, Wenbo Zhang et al. introduce RACER, which adaptively routes…
  7. 78 score ArXiv · Worth reading · 1 min CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation CADBench is a unified multimodal benchmark for recovering editable CAD programs from 2D/3D inputs, assembled from 18,000 samples across six families and five input modalities with six evaluation metrics. The authors evaluated 11 systems (>1.4M generated programs) and found mesh-to-CAD specialists beat code-generating VLMs under ideal inputs; they…
  8. 76 score ArXiv · Worth reading · 1 min Personal Visual Context Learning in Large Multimodal Models Personal Visual Context Learning (Personal VCL) formalizes how wearable-driven LMMs should use user-specific first-person visual evidence to resolve personalized queries. Xue et al. introduce Personal-VCL-Bench (covers persons, objects, behaviors), identify a pronounced context-utilization gap in frontier LMMs, and propose the Agentic Context…
  9. 72 score ArXiv · Worth reading · 1 min Factual recall in linear associative memories: sharp asymptotics and mechanistic insights The paper analyzes limits of factual recall in linear associative memories by introducing a decoupled model and using statistical‑physics tools to obtain sharp asymptotics. It proves equivalence between the decoupled and original models (capacity, weight spectra, mechanism) and derives the capacity law p_c log p_c / d^2 = 1/2. The work explains…
  10. 70 score ArXiv · Worth reading · 1 min Count Anything at Any Granularity Count Anything at Any Granularity (Liu et al., 2026) tackles brittle open-world object counting by making counting granularity explicit across five levels (identity, attribute, instance type, category, abstract concept). The authors build KubriCount using a novel automatic pipeline combining controllable 3D synthesis, image editing, and VLM…
  11. 68 score Twitter/X · Worth reading · 2 min Associated Builders and Contractors projects the industry needs 349,000 new… Associated Builders and Contractors projects the industry needs 349,000 new workers in 2026 and 456,000 in 2027; 41% of the current construction workforce will retire by 2031, while specialty-trade wages are rising 4–11% and some firms pay 20% premiums to retain crews.
  12. 65 score ArXiv · Worth reading · 1 min Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with $k$-step Policy Gradients Revisiting policy gradients for restricted policy classes, the paper identifies one-step myopia (policy gradients depending only on the one-step Q) as a cause of suboptimal critical points and proposes a k-step policy gradient that couples randomness across k steps. The method yields performance exponentially close in k to the optimal…
  13. 62 score ArXiv · Worth reading · 1 min A Recursive Decomposition Framework for Causal Structure Learning in the Presence of Latent Variables The paper introduces DiCoLa, a recursive decomposition framework that extends divide-and-conquer causal discovery to settings with latent variables by splitting the global CI-testing task into…
  14. 62 score substack.com · Worth reading · 14 min Introducing Barely Possible Barely Possible is Lane Rettig’s AI-operated podcast, launched publicly on May 12, 2026 (BarelyPossible.to and on Transistor). Rettig positions the show as a high-quality, fully automated daily…
  15. 62 score ArXiv · Worth reading · 1 min Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping Text-to-image post-training via reinforcement learning faces reward-hacking and miscalibration from normalization. Based on the abstract, the authors propose Super-Linear Advantage Shaping (SLAS)…
  16. 38 score ArXiv · Quick skim · 1 min Is Your Driving World Model an All-Around Player? WorldLens introduces a comprehensive benchmark for driving world models, spanning five aspects and…
  17. 35 score ArXiv · Quick skim · 1 min Safe Aerial 3D Path Planning for Autonomous UAVs using Magnetic Potential Fields 3DMaxConvNet extends the MaxConvNet magnetic potential-field planner into 3D by training a…
  18. 35 score ArXiv · Quick skim · 1 min When Can Digital Personas Reliably Approximate Human Survey Findings? LLM-based digital personas were evaluated as substitutes for human survey respondents using the…
  19. 35 score producthabits.com · Quick skim · 1 min Half of you got a test email yesterday Hiten Shah (Product Habits) ran an A/B email test on 2026-05-12 — half of subscribers saw one…
  20. 35 score Twitter/X · Quick skim · 1 min OpenSquilla: @hasantoxr ran a 10+ turn research workflow (published 2026-05-12)… OpenSquilla: @hasantoxr ran a 10+ turn research workflow (published 2026-05-12) and says smart…
  21. 35 score ArXiv · Quick skim · 1 min PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models PriorVLA introduces a prior-preserving adaptation method for Vision-Language-Action models that…
  22. 35 score ArXiv · Quick skim · 1 min Optimal and Scalable MAPF via Multi-Marginal Optimal Transport and Schrödinger Bridges The paper addresses anonymous MAPF by formulating it as a Markovian MMOT that exactly reduces an…
  23. 35 score ArXiv · Quick skim · 1 min RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark RoboMemArena is a large-scale benchmark addressing robotic memory shortcomings by offering 26…
  24. 35 score ArXiv · Quick skim · 1 min Unified Noise Steering for Efficient Human-Guided VLA Adaptation UniSteer tackles efficient real‑world adaptation of diffusion-based vision‑language‑action models…
  25. 35 score ArXiv · Quick skim · 1 min Counterfactual Stress Testing for Image Classification Models Counterfactual stress testing for medical image classifiers uses causal generative models to create…
  26. 35 score ArXiv · Quick skim · 1 min Confidence-Guided Diffusion Augmentation for Enhanced Bangla Compound Character Recognition Handwritten Bangla compound character recognition is improved via a confidence-guided diffusion…
  27. 35 score ArXiv · Quick skim · 1 min HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models HarmoWAM presents an end-to-end World Action Model that unifies predictive and reactive control by…
  28. 35 score Twitter/X · Quick skim · 1 min @KayvonJafar cites Andrej Karpathy’s line “I don’t think I’ve typed a line of… @KayvonJafar cites Andrej Karpathy’s line “I don’t think I’ve typed a line of code since December”…
  29. 35 score ArXiv · Quick skim · 1 min Price of Quality: Sufficient Conditions for Sparse Recovery using Mixed-Quality Data Sparse recovery from mixed-quality measurements (n1 high-quality, n2 low-quality) is studied via…
  30. 35 score ArXiv · Quick skim · 1 min Pixal3D: Pixel-Aligned 3D Generation from Images Pixal3D (Li et al., SIGGRAPH 2026) presents a pixel-aligned 3D generation paradigm that…
  31. 35 score Twitter/X · Quick skim · 1 min Profound's Marketing Engineering Hackathon (June 6, 2026, Union Square, NYC) will… Profound's Marketing Engineering Hackathon (June 6, 2026, Union Square, NYC) will host 50 builders…
  32. 35 score ArXiv · Quick skim · 1 min Variational Inference for Lévy Process-Driven SDEs via Neural Tilting Variational Inference for Lévy Process-Driven SDEs via Neural Tilting introduces a neural…
  33. 35 score ArXiv · Quick skim · 1 min Geometry-aware Prototype Learning for Cross-domain Few-shot Medical Image Segmentation GeoProto addresses cross-domain few-shot medical image segmentation by enriching prototypical…
  34. 35 score Twitter/X · Quick skim · 1 min When you trade on Polymarket with PolyHelper, you automatically farm two airdrops… When you trade on Polymarket with PolyHelper, you automatically farm two airdrops — Polymarket’s…
  35. 35 score ArXiv · Quick skim · 1 min When Are Trade-Off Functions Testable from Finite Samples? The paper studies finite-sample hypothesis testing for the type I/type II error trade-off function…
  36. 35 score Twitter/X · Quick skim · 1 min CloakBrowser is a MIT‑licensed stealth Chromium (pip install, ~200MB) that the… CloakBrowser is a MIT‑licensed stealth Chromium (pip install, ~200MB) that the author says achieves…
  37. 35 score ArXiv · Quick skim · 1 min CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models CapVector introduces a method to decouple auxiliary-objective finetuning into two parameter-space…
datacenterdynamics.com · 1 min Signal

Stream Now: DCD>Inside APAC

DCD Investment & Markets Channel's livestream “DCD>Inside APAC” (published 11 May 2026) analyzes how rapid AI investment and sovereign cloud mandates are reshaping where and how data centres are built across Singapore, Sydney, Tokyo and emerging markets in Indonesia, Malaysia and the Philippines. Key focuses include grid and land constraints, government‑industry sustainability alignment, novel deal structures to unlock capacity, and technical approaches like edge deployments and cross‑border interconnection.

Open reader
Must read

Start here. These are the items with the strongest reader value today.

3 items
2 datacenterdynamics.com 2026-05-12 2 min read
Open

Your streaming link: Keeping IT Cool: Liquid 2026 Day Two

Keeping IT Cool | Liquid — Day Two of DCD's Liquid series streams live at 9am ET (2pm BST) on Wednesday, 13 May 2026, focusing on liquid cooling for AI-ready data centres. Sessions cover >100 kW rack densities, scalable cooling loops, retrofit strategies, chiller-to-chip architectures, dry-break quick-disconnect reliability testing, and research from Georgia Tech, Binghamton and IBM.

By James Raddings
3 datacenterdynamics.com 2026-05-12 1 min read
Open

What’s driving India’s next wave of data center expansion?

NTT Global Data Centers' whitepaper (Data Centre Dynamics, 12 May 2026) frames India as an emerging global digital‑infrastructure hub, driven by AI workloads, fast cloud adoption, data‑localization rules and growing enterprise demand. It analyzes market dynamics, infrastructure trends, investment drivers and regulatory/operational issues to help operators plan long‑term capacity, connectivity and cloud expansion.

By DCD Cloud & Hybrid
Worth reading

Useful context and follow-up reading when you have more time.

12 items
1 divenewsletter.com 2026-05-12 1 min read
Open

Watch Webinar | Agentic AI for Holistic Major Event Management

Agentic AI for Holistic Major Event Management (Schneider Electric, Utility Dive, 2026-05-12) presents an integrated platform approach that unifies planning, real-time operations, and post-event recovery. It emphasizes the 5–7-day pre-event window and uses AI-augmented insights, wildfire modeling, vegetation analytics, and asset-health data to produce actionable, auditable, agent-driven plans that scale across the grid lifecycle.

By Utility Dive
2 ArXiv 2026-05-11 1 min read
Open

What should post-training optimize? A test-time scaling law perspective

Large-language-model post-training often optimizes single-response mean reward, creating a mismatch with best-of-N test-time selection. The paper analyzes the budget-mismatch setting (m << N) and proves that, assuming structure in reward tails, best-of-N policy gradients can be extrapolated from small-rollout samples. It introduces Tail-Extrapolated estimators (TEA, Prefix-TEA) and shows improved best-of-N performance on instruction-following benchmarks. Summary based on the abstract only.

Authors: Muheng Li, Jian Qian, Wenlong Mou
3 ArXiv 2026-05-11 1 min read
Open

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge

The paper studies LLM-as-a-Judge trade-offs and finds that explicit reasoning markedly boosts accuracy on structured verification tasks (e.g., math and coding) while offering limited or negative benefit on simpler evaluations and costing substantially more compute. To address this, Wenbo Zhang et al. introduce RACER, which adaptively routes examples between reasoning and non-reasoning judges under a fixed budget via distributionally robust optimization with a KL uncertainty set. RACER uses an efficient primal–dual solver, has theoretical guarantees (unique optimal policy, linear convergence), and empirically attains better accuracy–cost trade-offs under distribution shift. (Based on the paper abstract; ICML 2026 acceptance.)

Authors: Wenbo Zhang, Lijinghua Zhang, Liner Xiang...
4 ArXiv 2026-05-11 1 min read
Open

CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

CADBench is a unified multimodal benchmark for recovering editable CAD programs from 2D/3D inputs, assembled from 18,000 samples across six families and five input modalities with six evaluation metrics. The authors evaluated 11 systems (>1.4M generated programs) and found mesh-to-CAD specialists beat code-generating VLMs under ideal inputs; they report brittleness under modality shift and complexity-related failures. (Summary based on abstract; full text not available.)

Authors: Anna C. Doris, Jacob Thomas Sony, Ghadi Nehme...
5 ArXiv 2026-05-11 1 min read
Open

Personal Visual Context Learning in Large Multimodal Models

Personal Visual Context Learning (Personal VCL) formalizes how wearable-driven LMMs should use user-specific first-person visual evidence to resolve personalized queries. Xue et al. introduce Personal-VCL-Bench (covers persons, objects, behaviors), identify a pronounced context-utilization gap in frontier LMMs, and propose the Agentic Context Bank—a self‑refining memory plus query‑adaptive evidence selection—that consistently improves performance over standard context prompting across tasks and evaluated backbones.

Authors: Zihui Xue, Ami Baid, Sangho Kim...
6 ArXiv 2026-05-11 1 min read
Open

Factual recall in linear associative memories: sharp asymptotics and mechanistic insights

The paper analyzes limits of factual recall in linear associative memories by introducing a decoupled model and using statistical‑physics tools to obtain sharp asymptotics. It proves equivalence between the decoupled and original models (capacity, weight spectra, mechanism) and derives the capacity law pc log pc / d^2 = 1/2. The work explains why optimal learning nudges correct scores just above extreme‑value thresholds and generalizes results to two‑layer linear architectures, providing a baseline for memory in neural networks.

Authors: Alessio Giorlandino, Sebastian Goldt, Antoine Maillard
7 ArXiv 2026-05-11 1 min read
Open

Count Anything at Any Granularity

Count Anything at Any Granularity (Liu et al., 2026) tackles brittle open-world object counting by making counting granularity explicit across five levels (identity, attribute, instance type, category, abstract concept). The authors build KubriCount using a novel automatic pipeline combining controllable 3D synthesis, image editing, and VLM filtering, and find existing multimodal LLMs and counting models fail to follow fine-grained prompts. They train HieraCount, which fuses text and visual exemplars, and report substantial accuracy and generalization gains on multi-grained counting tasks.

Authors: Chang Liu, Haoning Wu, Weidi Xie
8 Twitter/X 2026-05-12 2 min read
Open

Associated Builders and Contractors projects the industry needs 349,000 new…

Why it matters

Associated Builders and Contractors projects the industry needs 349,000 new workers in 2026 and 456,000 in 2027; 41% of the current construction workforce will retire by 2031, while specialty-trade wages are rising 4–11% and some firms pay 20% premiums to retain crews.

  • Infill is the only practical supply path in desirable metros (SF, NYC, LA, Boston, Seattle, DC, Austin), but mid‑rise/tower infill costs 2–3× per sq ft and depends on hard‑to‑staff, union trades (ironworkers, concrete crews, complex MEP, elevator mechanics).
  • Austin permitted the most multifamily per capita from 2021–2023 and saw rents fall ~20% from the 2022 peak, but those units were built under the old code on cheap greenfield land with an ~40% immigrant workforce; Durham argues zoning alone won’t scale supply and advocates construction robotics and large prefabricated components to cut labor up to ~90%.

Nick Durham argues zoning wins won’t deliver housing quickly because the U.S. construction labor force is aging and undersized (ABC projects 349k hires in 2026, 456k in 2027; 41% retiring by 2031) and greenfield land and immigrant labor are drying up. He cites Austin’s 2021–2023 multifamily boom and 20% rent drop as driven by old‑code, low‑rise, edge‑land development, and pushes construction robotics and large prefabrication to sharply boost productivity.

By @americanhousing
9 ArXiv 2026-05-11 1 min read
Open

Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with $k$-step Policy Gradients

Revisiting policy gradients for restricted policy classes, the paper identifies one-step myopia (policy gradients depending only on the one-step Q) as a cause of suboptimal critical points and proposes a k-step policy gradient that couples randomness across k steps. The method yields performance exponentially close in k to the optimal deterministic policy, converges in O(1/T) with projected/mirror descent under mild smoothness/differentiability assumptions, avoids common distribution-mismatch factors, and targets applications like state aggregation and partially observable cooperative multi-agent problems. Full text on arXiv.

Authors: Alex DeWeese, Guannan Qu
10 ArXiv 2026-05-11 1 min read
Open

A Recursive Decomposition Framework for Causal Structure Learning in the Presence of Latent Variables

The paper introduces DiCoLa, a recursive decomposition framework that extends divide-and-conquer causal discovery to settings with latent variables by splitting the global CI-testing task into smaller subproblems and integrating solutions via a principled reconstruction step. The authors prove soundness and completeness for the framework and demonstrate substantial runtime improvements on synthetic experiments and practical effectiveness on a real-world dataset, addressing high-dimensional CI-testing bottlenecks.

Authors: Zheng Li, Feng Xie, Shenglan Nie...
11 substack.com 2026-05-12 14 min read
Open

Introducing Barely Possible

Barely Possible is Lane Rettig’s AI-operated podcast, launched publicly on May 12, 2026 (BarelyPossible.to and on Transistor). Rettig positions the show as a high-quality, fully automated daily briefing—initially a single AI-produced program targeted at roughly one to two hours per day that synthesizes world news, AI developments, crypto and tech/finance updates, and a focused daily deep dive. He frames the product around real listening habits (he reports ~1–2 hours/day of audio and uses 2x speed to get through 3–4 episodes on a run) and wants a compact, high-octane slot that replaces time otherwise spent hunting for disparate sources.

Technically, Barely Possible runs a content-collection-and-analysis engine that synthesizes consistent underlying stories into different user-specific outputs; Rettig says the pipeline can operate on local open models, keeping costs sustainable. The project began about two months before launch with collaborator “Baz,” and Rettig has iterated on it daily. The broader vision is an AI-native audio platform with effectively infinite, customizable content (per-user “transcoding” of a single source into many tailored formats), personalized advertising, and creator tools—drawing on concepts like skill files and Neal Stephenson’s ‘ractive’—rather than attempting to supplant human-hosted long-form podcasts.

By Lane Rettig from Three Things
12 ArXiv 2026-05-11 1 min read
Open

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping

Text-to-image post-training via reinforcement learning faces reward-hacking and miscalibration from normalization. Based on the abstract, the authors propose Super-Linear Advantage Shaping (SLAS), which extends the Fisher–Rao metric with advantage-dependent weighting to amplify informative updates and suppress illusory gradients, plus batch-level normalization. Evaluations report consistent gains over DanceGRPO, faster training, better GenEval and UniGenBench++ OOD performance, and improved scaling robustness and fidelity.

Authors: Haoyuan Sun, Jing Wang, Yuxin Song...
Quick skim

Scan these for facts, links, or weak signals worth tracking.

22 items · open
1 ArXiv 2026-05-11 1 min read
Open

Is Your Driving World Model an All-Around Player?

WorldLens introduces a comprehensive benchmark for driving world models, spanning five aspects and 24 dimensions to evaluate pixels, 4D geometry, closed-loop driving, and human perceptual realism. The authors benchmark six models, show no method excels across axes (best attain 2–3/10 human realism), and provide WorldLens-26K plus WorldLens-Agent for scalable, explainable evaluation. (Summary based on abstract; full text not reviewed.)

Authors: Lingdong Kong, Ao Liang, Tianyi Yan...
2 ArXiv 2026-05-11 1 min read
Open

Safe Aerial 3D Path Planning for Autonomous UAVs using Magnetic Potential Fields

3DMaxConvNet extends the MaxConvNet magnetic potential-field planner into 3D by training a convolutional autoencoder to generate obstacle-aware potential fields from LiDAR-derived 101^3 voxel grids. In Cosys-AirSim experiments (100 randomized closed-loop trials on two urban maps) it reached 100% success without retraining, matched A* path quality with ~2× lower runtime, and ran ~200× faster than RRT*(3k).

Authors: Haechan Mark Bong, Giovanni Beltrame
3 ArXiv 2026-05-11 1 min read
Open

When Can Digital Personas Reliably Approximate Human Survey Findings?

LLM-based digital personas were evaluated as substitutes for human survey respondents using the LISS panel: personas were built from background variables and pre-2023 survey histories and tested on held-out post-cutoff answers. Across four persona architectures, three LLMs, and two prediction tasks, personas improved distributional alignment (notably for stable attributes) but struggled with individual prediction and multivariate respondent structure; retrieval augmentation helped. Summary is based on the abstract (full paper not reviewed).

Authors: Mumin Jia, Yilin Chen, Divya Sharma...
4 producthabits.com 2026-05-12 1 min read
Open

Half of you got a test email yesterday

Hiten Shah (Product Habits) ran an A/B email test on 2026-05-12 — half of subscribers saw one version and half another, with some calling one variant AI-sounding. Facing a year-long drop in click rates, Shah hypothesizes readers detect “empty” AI-style prose and will run a live Reddit AMA to probe responses and defend human writing.

By Hiten Shah
5 Twitter/X 2026-05-12 1 min read
Open

OpenSquilla: @hasantoxr ran a 10+ turn research workflow (published 2026-05-12)…

OpenSquilla: @hasantoxr ran a 10+ turn research workflow (published 2026-05-12) and says smart routing routed simple prompts to cheaper models and complex ones to full power, cutting his token bill by ~67% versus his usual agent while preserving memory and output quality. OpenSquilla (open-source, Python, local‑first) claims 60–80% cost reduction and launched the #10MTokenChallenge (30 winners × 10M OpenRouter credits).

By @hasantoxr
6 ArXiv 2026-05-11 1 min read
Open

PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models

PriorVLA introduces a prior-preserving adaptation method for Vision-Language-Action models that freezes a Prior Expert and trains an Adaptation Expert, using Expert Queries to inject pretrained scene and motor priors. By updating only 25% of parameters, it outperforms full fine-tuning and SOTA baselines on RoboTwin 2.0, LIBERO, and eight real-world tasks, with strong OOD and few-shot gains.

Authors: Xinyu Guo, Bin Xie, Wei Chai...
7 ArXiv 2026-05-11 1 min read
Open

Optimal and Scalable MAPF via Multi-Marginal Optimal Transport and Schrödinger Bridges

The paper addresses anonymous MAPF by formulating it as a Markovian MMOT that exactly reduces an exponentially large problem to a polynomial-size LP, with proven feasibility and total unimodularity yielding collision-free integral {0,1} transports. For large-scale instances the authors use a Schrödinger-bridge entropic regularization solved via Sinkhorn iterations to produce fractional templates used in a reduced LP, trading negligible optimality loss for major computational savings; experimental results demonstrate strong optimality and scalability. Accepted to ICML 2026 as a spotlight paper.

Authors: Usman A. Khan, Joseph W. Durham
8 ArXiv 2026-05-11 1 min read
Open

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark

RoboMemArena is a large-scale benchmark addressing robotic memory shortcomings by offering 26 long-horizon tasks (avg >1,000 steps) with multimodal memory annotations and paired real-world evaluations; a VLM-based pipeline composes subtasks and generates trajectories via atomic functions. The authors introduce PrediMem, a dual-system VLA with recent/keyframe memory buffers and a predictive coding head that outperforms baselines on this benchmark.

Authors: Huashuo Lei, Wenxuan Song, Huarui Zhang...
9 ArXiv 2026-05-11 1 min read
Open

Unified Noise Steering for Efficient Human-Guided VLA Adaptation

UniSteer tackles efficient real‑world adaptation of diffusion-based vision‑language‑action models by combining human corrective interventions with noise‑space RL. The method inverts a frozen flow‑matching decoder to convert human actions into supervision for a noise‑predicting actor, which is concurrently improved with RL. Real‑world tests on four manipulation tasks show a jump from 20% to 90% success in 66 minutes on average. Full paper was not available, summary based on the abstract.

Authors: Junjie Lu, Xinyao Qin, Yuhua Jiang...
10 ArXiv 2026-05-11 1 min read
Open

Counterfactual Stress Testing for Image Classification Models

Counterfactual stress testing for medical image classifiers uses causal generative models to create realistic, anatomically-preserving ‘‘what-if’’ images by intervening on attributes like scanner type and patient sex. Evaluated on chest X‑ray and mammography across three architectures and several shift scenarios (Stammel et al., 2026), the method more reliably predicts real OOD performance than simple brightness/contrast perturbations, improving robustness assessment prior to deployment.

Authors: Moritz Stammel, Fabio De Sousa Ribeiro, Raghav Mehta...
11 ArXiv 2026-05-11 1 min read
Open

Confidence-Guided Diffusion Augmentation for Enhanced Bangla Compound Character Recognition

Handwritten Bangla compound character recognition is improved via a confidence-guided diffusion augmentation approach that synthesizes class-conditional samples, enhances U-Net blocks with Squeeze-and-Excitation residuals, and filters generated images using pre-trained classifiers. Fusing filtered synthetic images with real data yields consistent gains across ResNet50, DenseNet121, VGG16 and ViT, with a top accuracy of 89.2% on AIBangla. Only the abstract was available for this summary.

Authors: Md. Sultan Al Rayhan, Maheen Islam
12 ArXiv 2026-05-11 1 min read
Open

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

HarmoWAM presents an end-to-end World Action Model that unifies predictive and reactive control by conditioning a predictive expert and a reactive expert on spatio-temporal priors from a world model, with a Process-Adaptive Gating Mechanism to coordinate switching. Reported zero-shot generalization covers three unseen real-world environments and six manipulation tasks, improving over VLA and WAM baselines by 33% and 29%, respectively. Full paper text not available (abstract only).

Authors: Qiuxuan Feng, Jiale Yu, Jiaming Liu...
13 Twitter/X 2026-05-11 1 min read
Open

@KayvonJafar cites Andrej Karpathy’s line “I don’t think I’ve typed a line of…

@KayvonJafar cites Andrej Karpathy’s line “I don’t think I’ve typed a line of code since December” and Garry Tan’s design prompt to have one person run a whole software team. He points to the open-source gstack as that answer, with Garry claiming ~810x higher pace; gstack wires AI roles (CEO, eng, QA, browser operator) into workflows like /office-hours, /autoplan, /qa, /review, pair-agent and parallel sprints, reframing development as operating systems of AI workers.

By @KayvonJafar
14 ArXiv 2026-05-11 1 min read
Open

Price of Quality: Sufficient Conditions for Sparse Recovery using Mixed-Quality Data

Sparse recovery from mixed-quality measurements (n1 high-quality, n2 low-quality) is studied via information-theoretic and algorithmic sample-size conditions. The paper proves a linear 'Price of Quality' trade-off: agnostic decoders cap one high-quality sample at no more than two low-quality ones, whereas informed decoders can exploit variance knowledge to make the price unbounded. LASSO matches homogeneous-noise thresholds, depending only on average noise. Published ICLR 2026 (arXiv 2026-05-11).

Authors: Youssef Chaabouni, David Gamarnik
15 ArXiv 2026-05-11 1 min read
Open

Pixal3D: Pixel-Aligned 3D Generation from Images

Pixal3D (Li et al., SIGGRAPH 2026) presents a pixel-aligned 3D generation paradigm that back-projects multi-scale image features into a 3D feature volume to create explicit pixel-to-3D correspondences, generating assets aligned with the input view. The method reportedly substantially raises fidelity—'approaching the fidelity level of reconstruction'—and extends to multi-view and scene synthesis.

Authors: Dong-Yang Li, Wang Zhao, Yuxin Chen...
16 Twitter/X 2026-05-12 1 min read
Open

Profound's Marketing Engineering Hackathon (June 6, 2026, Union Square, NYC) will…

Profound's Marketing Engineering Hackathon (June 6, 2026, Union Square, NYC) will host 50 builders for a one-day challenge to automate inhuman-scale marketing processes, offering $40,000 total in prizes and judges from Ramp, Stripe, and MongoDB. The platform-agnostic event allows Claude Code, Cursor, n8n, Python, LangChain, or raw APIs; winners receive cash, Profound credits, and interviews. Application required.

By @hasantoxr
17 ArXiv 2026-05-11 1 min read
Open

Variational Inference for Lévy Process-Driven SDEs via Neural Tilting

Variational Inference for Lévy Process-Driven SDEs via Neural Tilting introduces a neural exponential tilting framework that reweights the Lévy measure with neural nets to build a flexible variational family preserving jumps. A quadratic parametrization gives closed-form normalization; a conditional Gaussian representation for stable processes enables simulation and scalable, symmetry-aware Monte Carlo optimization. Summary based on the abstract (full text not reviewed).

Authors: Yaman Kindap, Manfred Opper, Benjamin Dupuis...
18 ArXiv 2026-05-11 1 min read
Open

Geometry-aware Prototype Learning for Cross-domain Few-shot Medical Image Segmentation

GeoProto addresses cross-domain few-shot medical image segmentation by enriching prototypical matching with explicit geometric priors. The method's core, GAPE, augments appearance prototypes with learned ordinal offsets from an Ordinal Shape Branch trained under an ordinal-consistency loss (no extra labels beyond segmentation masks). According to the abstract, GeoProto attains state-of-the-art results across seven datasets and three settings (cross-modality, cross-sequence, cross-context). Only the paper abstract was available for this briefing.

Authors: Feifan Song, Yuntian Bo, Haofeng Zhang
19 Twitter/X 2026-05-12 1 min read
Open

When you trade on Polymarket with PolyHelper, you automatically farm two airdrops…

When you trade on Polymarket with PolyHelper, you automatically farm two airdrops — Polymarket’s and a PolyHelper airdrop — while PolyHelper reimburses trading fees (they refunded all fees paid to Polymarket on a randomly chosen day, May 8). PolyHelper also adds extra UX info/features and is free; the author (@0xd1namit) urges there’s no reason not to use it.

By @0xd1namit
20 ArXiv 2026-05-11 1 min read
Open

When Are Trade-Off Functions Testable from Finite Samples?

The paper studies finite-sample hypothesis testing for the type I/type II error trade-off function between two unknown distributions. Without structure testing is impossible; they identify exact attainability of Neyman–Pearson rejection regions by a class S and show finite VC-dimension of S is necessary and sufficient. They give a test with nonasymptotic type I control and uniform power under explicit separations, produce simultaneous confidence bands, analyze local rates (monotone likelihood-ratio) with near-matching lower bounds, and extend to approximate attainability (e.g., univariate log-concave).

Authors: Kaining Shi, Qiaosen Wang, Cong Ma
21 Twitter/X 2026-05-12 1 min read
Open

CloakBrowser is a MIT‑licensed stealth Chromium (pip install, ~200MB) that the…

CloakBrowser is a MIT‑licensed stealth Chromium (pip install, ~200MB) that the author says achieves human-like anti-bot fidelity — 0.9 on reCAPTCHA v3 versus stock Playwright's 0.1 — by patching Chromium's C++ in 16 spots (canvas, WebGL, audio, fonts, hardwareConcurrency, GPU strings, WebDriver flag, TLS fingerprint). It claims to pass 14/14 detection tests, auto-resolve Turnstile, and replace commercial $200–$500+/month stealth browsers or costly custom builds.

By @hasantoxr
22 ArXiv 2026-05-11 1 min read
Open

CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models

CapVector introduces a method to decouple auxiliary-objective finetuning into two parameter-space components: general capability enhancement and task-specific action fitting. By finetuning two small-scale models with different strategies and taking their parameter difference as a capability vector, then merging it into pretrained VLA models (plus an orthogonal regularizer), standard SFT attains auxiliary-like gains with lower compute. Full text not provided here.

Authors: Wenxuan Song, Han Zhao, Fuhao Li...