Briefing · 2026-06-27

Your briefing

64 ranked ·

Today's dispatch

Filed · 64 ranked

  1. 84 score ArXiv · Worth reading · 1 min Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models VISE (Visual Invariance Self-Evolution) is a purely unsupervised, single-model self-evolving LMM framework that targets 'visual under-conditioning' by directly regularizing visual conditioning with two invariance-based rewards: a geometric invariance reward (spatial consistency under known transformations) and a semantic invariance reward (penalizes evidence-agnostic generation when predicted regions are perturbed).
  2. 82 score ArXiv · Worth reading · 1 min The Geometry of Updates: Fisher Alignment at Vocabulary Scale Representation metrics (e.g., CKA) are non-identifiable for transfer in shared-output-head settings: models can share identical activations yet have orthogonal head updates. The paper proves an identity that head Fisher alignment equals the cosine between kernel mean embeddings in the joint activation–error space.
  3. 80 score ArXiv · Worth reading · 1 min LMs as Task-Specific Knowledge Bases: An Interpretability Analysis Elhelo, Globerson, and Geva (ArXiv preprint, 2026-06-25) show language models encode factual knowledge in a task-specific way: behavioral experiments find facts acquired on one task frequently fail to co-emerge on others during training.
  4. 78 score Twitter/X · Worth reading · 2 min Frontier models concentrate economic value Frontier models concentrate economic value: customers currently pay for Anthropic and OpenAI while firms like Mistral and XAI are falling behind, producing winner-take-most dynamics and (in the most valuable markets) potentially winner-take-all outcomes.
  5. 76 score ArXiv · Worth reading · 1 min Any system that outputs one member model's answer is upper-bounded by accuracy ≤… Any system that outputs one member model's answer is upper-bounded by accuracy ≤ 1 − β, where β is the rate that every model is wrong on the same query; a Clopper–Pearson bound on β gives a finite-sample certificate on the maximum possible gain from routing, voting, or cascades.
  6. 76 score ArXiv · Worth reading · 1 min Reinforcement Learning without Ground-Truth Solutions can Improve LLMs RiVER (Ranking-induced VERifiable) trains LLMs using deterministic execution feedback as continuous rewards on score-based tasks, addressing 'scale dominance' and 'frequency dominance' via instance-wise calibrated reward shaping that emphasizes top-ranked solvers while keeping bounded feedback for other valid solutions; training used 12 AtCoder Heuristic Contest tasks.
  7. 75 score ArXiv · Worth reading · 1 min Hallucination in World Models is Predictable and Preventable MMBench2: a 427-hour, 210-task visual world-modeling dataset with ground-truth actions, rewards, and live simulators; authors train a 350M-parameter generative world model on it.
  8. 72 score ArXiv · Worth reading · 1 min Multilingual Reasoning Cascades Need More Context A simple, training-free context-aware cascade—giving the final translation module the original question, the English-translated question, and the English reasoning trace—yields strong gains for open-ended generation across nine multilingual benchmarks, three backbone models, and 285 high-/mid-/low-resource languages (ArXiv, 2026-06-25).
  9. 72 score ArXiv · Worth reading · 1 min Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance Proposes a training-free, plug-and-play feature self-guidance that disperses internal features during batch generation and uses a manifold regularization step to project dispersed features back onto the data manifold, aiming to mitigate diversity collapse in pretrained flow models.
  10. 72 score ArXiv · Worth reading · 1 min VibeAct: Vibration to Actions for Contact-Rich Reactive Robot Dexterity VibeAct embeds piezoelectric microphones in a dexterous robot hand, records vibro-acoustic data via teleoperation, and replays recordings in a calibrated digital clone to auto-label per-finger contact and slip for training a tactile estimator.
  11. 72 score ArXiv · Worth reading · 1 min How Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple Mitigation The paper introduces a four-dimension diagnostic (tokenization cost, predictive uncertainty/surprisal, semantic robustness, context sensitivity) and evaluates it on three corpora: 17th‑century Italian (1610–1689), 19th‑century I Promessi Sposi (high-exposure control), and 18th‑century Russian civil print books.
  12. 70 score ArXiv · Worth reading · 1 min Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders Top-k SAE augmented with two sparsity regularizers—an ℓ1 penalty on off-support (unselected) units and a scale-invariant ℓ1/ℓ2 ratio on batch-active units—by Jacquier, Vakalopoulou, and Hosseini (published 2026-06-25) increases monosemanticity across two datasets, three vision foundation models, and a range of k while preserving reconstruction quality.
  13. 68 score ArXiv · Worth reading · 1 min Decision-Aligned Evaluation of Uncertainty Quantification Authors Annika Schneider, Tommy Rochussen, Joshua Stiller, and Vincent Fortuin (ArXiv:2606.26990v1, published 2026-06-25) introduce a decision-alignment criterion and show that common UQ metrics—negative log-likelihood (NLL) and expected calibration error (ECE)—can be misaligned with downstream decision utility.
  14. 68 score ArXiv · Worth reading · 1 min Scalable Behavior Cloning with Open Data, Training, and Evaluation ABC-130K: released the largest open-source teleoperation dataset to date with 3,500 hours across ~130K episodes covering 195 manipulation tasks (paper posted 2026-06-25).
  15. 68 score ArXiv · Worth reading · 1 min When are likely answers right? On Sequence Probability and Correctness in LLMs Across prompt–answer pairs within a fixed dataset, higher sequence probability often predicts correctness: Zenn & Geiping (2026) quantify this alignment at four levels (decoding methods, hyperparameters, prompt–answer pairs, repeated responses) in a 38‑page preprint (10 pages main, 28 pages appendix).
  16. 68 score ArXiv · Worth reading · 1 min Error-Conditioned Neural Solvers Error-Conditioned Neural Solvers (ENS) (Jiang et al., arXiv 2026-06-25) feed the PDE residual field into the network at every iteration so the model learns an update policy to correct its own errors, avoiding expensive classical hybrid optimizers (gradient descent / Gauss–Newton) and their instability.
  17. 68 score ArXiv · Worth reading · 1 min Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards Proposes a self-evolving training framework (published 2026-06-25) that decomposes a unified LMM into three internal roles—Proposer (generates visual questions), Solver (answers and evaluates), and Generator (synthesizes images)—and trains using only unlabeled images with no human annotations, preference labels, or external reward/judge models.
  18. 68 score ArXiv · Worth reading · 1 min Bridging Performance and Generalization in Reinforcement Learning for Agile Flight The authors combine task-aware switching (based on learning progress) with a physically informed procedural track generator to produce a fast, robust zero-shot generalist policy for drone racing, achieving a 7.4x improvement in generalization over prior state-of-the-art while keeping competitive racing speeds and requiring no test-time adaptation.
  19. 62 score Twitter/X · Worth reading · 1 min cryptopunk7213 claims every AI model will fall into two groups — 'Brain' (highly… cryptopunk7213 claims every AI model will fall into two groups — 'Brain' (highly capable) and 'Executor' (workers) — and whoever builds the smart‑router platform on top will capture outsized value; 'Brain' models will cost ~5–10× more per token and Fortune 100 companies will pay for them.
  20. 62 score ArXiv · Worth reading · 1 min Ribbon: Scalable Approximation and Robust Uncertainty Quantification Ribbon (Gibson, Tipton, Rumsey, Klein; arXiv 2026-06-25) approximates Dirichlet-reweighted (Bayesian/weighted-likelihood) bootstrap by using an influence-function linearization around a single fitted model, replacing repeated refitting with post-hoc linear algebra.
  21. 62 score ArXiv · Worth reading · 1 min Hamid Reza Firoozfar et al. Hamid Reza Firoozfar et al. (published 2026-06-25; arXiv:2606.27314v1) introduce a comprehensive, mechanism-oriented taxonomy of indirect linguistic expressions (ILE) that categorizes the underlying encoding and recovery operations rather than communicative goals.
  22. 62 score Twitter/X · Worth reading · 1 min GPT‑5.6 will run on Cerebras hardware at up to 750 tokens per second; OpenAI says… GPT‑5.6 will run on Cerebras hardware at up to 750 tokens per second; OpenAI says GPT‑5.6 Sol on Cerebras will be available in July 2026 with initial access limited to select customers.
  23. 60 score ArXiv · Worth reading · 1 min Continual Robot Policy Learning via Variational Neural Dynamics Proposes a continual learning framework that combines an analytical physics prior with a neural residual and a recurrent encoder that infers hidden, recurring dynamics from recent state-action trajectories; the inferred latent conditions condition both the residual dynamics model and the policy, and policy learning is done via differentiable simulation over sampled latent dynamics.
  24. 58 score ArXiv · Worth reading · 1 min OctoSense: Self-Supervised Learning for Multimodal Robot Perception OctoSense releases a 59-hour time-synchronized driving dataset (published 2026-06-25) containing stereo RGB + event cameras, LiDAR, thermal camera, IMU, RTK GPS, and proprioception (CAN bus and quadruped joint angles) across varied environments and times of day, including highly degraded-sensor scenarios.
  25. 58 score Twitter/X · Worth reading · 1 min Ornith-1.0 is an open-source family of agentic-coding LLMs with model sizes 9B… Ornith-1.0 is an open-source family of agentic-coding LLMs with model sizes 9B Dense, 31B Dense, 35B MoE, and 397B MoE; all models are released under the MIT license and were post-trained on top of gemma4 and qwen3.5.
  26. 42 score Twitter/X · Quick skim · 1 min On 2026-06-26, @johnloeber argues that in rivalries like the US vs China in AI… On 2026-06-26, @johnloeber argues that in rivalries like the US vs China in AI, when one side slows the other 'goes even faster'—pushing harder to overtake or to establish a permanent lead.
  27. 35 score ArXiv · Quick skim · 1 min Language-Based Digital Twins for Elderly Cognitive Assistance Proposes a language-based digital twin framework that leverages large language models plus stylometric cues and contextual metadata, and introduces a multi-head conditional variational autoencoder (cVAE) to jointly measure reconstruction quality and predict cognitive (MoCA) scores.
  28. 35 score Twitter/X · Quick skim · 1 min Mitchell Hashimoto announced on 2026-06-23 that his family is donating another… Mitchell Hashimoto announced on 2026-06-23 that his family is donating another $400,000 to the Zig Software Foundation.
  29. 35 score Twitter/X · Quick skim · 1 min @joshrauh asserts Gavin Newsom's premise that the wealthy pay lower tax rates… @joshrauh asserts Gavin Newsom's premise that the wealthy pay lower tax rates than the rest 'is a lie,' citing Auten and Splinter showing the average tax rate rises from 2% for the lowest quintile to 45% for the top 0.01% of taxpayers.
  30. 35 score Twitter/X · Quick skim · 1 min swyx says they’ve been “scaling without slop” by working with aligned domain… swyx says they’ve been “scaling without slop” by working with aligned domain experts to add coverage and co-organized their first AI Forward Deployed Engineering (FDE) miniconference with Basil Chatha at AI Engineer World’s Fair.
  31. 35 score Twitter/X · Quick skim · 1 min On June 25, 2026 Governor Gavin Newsom announced California’s… On June 25, 2026 Governor Gavin Newsom announced California’s “first-in-the-nation” dashboard to proactively track AI-related job-loss trends; the tool was developed with the University of California and the California Policy Lab and released under his recent executive order (GovPressOffice tweet).
  32. 35 score Twitter/X · Quick skim · 1 min At age 20, Turan Selvi says he was living with his parents in a small village in… At age 20, Turan Selvi says he was living with his parents in a small village in Germany, peaked at 310 pounds, was depressed and barely sleeping.
  33. 35 score Twitter/X · Quick skim · 3 min Tony's 'human in the loop = 20x pricing' play Tony's 'human in the loop = 20x pricing' play: a buddy's end‑to‑end AI A/B‑testing and funnel product got no traction as $500/mo software but, repackaged identically as a human‑accountable service, started closing $15,000/month contracts.
  34. 35 score Twitter/X · Quick skim · 2 min Gina Acosta (@ginacostag_) published a "Top 20 Free AI Tools You Need in 2026"… Gina Acosta (@ginacostag_) published a "Top 20 Free AI Tools You Need in 2026" list on 2026-06-26, categorizing tools across Productivity, Writing, Audio, App/Web dev, Coding, and Image/Design and naming examples like GetStudyPal, Perplexity, Google Gemini, ChatGPT, Grammarly, ElevenLabs, Replit, GitHub Copilot, MidJourney, Canva Magic Studio, and Ideogram.
  35. 35 score Twitter/X · Quick skim · 1 min In 1999 Pontevedra's mayor removed street parking, eliminated most… In 1999 Pontevedra's mayor removed street parking, eliminated most through‑traffic, and restricted car access to residents, deliveries, taxis and emergencies.
  36. 35 score Twitter/X · Quick skim · 1 min Nori's SuperNori is presented as the first AI that 'actually runs a household' —… Nori's SuperNori is presented as the first AI that 'actually runs a household' — it auto-restocks groceries, books travel, plans weekly meals around family nutrition goals, and controls smart-home devices, acting proactively without being asked.
  37. 35 score Twitter/X · Quick skim · 1 min Ben Horowitz (a16z) said in a January show appearance that California's proposed… Ben Horowitz (a16z) said in a January show appearance that California's proposed wealth tax is "the best strategy" he's seen to dismantle Silicon Valley's network effect.
  38. 35 score Twitter/X · Quick skim · 1 min @iruletheworldmo (posted 2026-06-26) states the era of public access to… @iruletheworldmo (posted 2026-06-26) states the era of public access to bleeding‑edge models is over, warning that 'if these models are being taken away, we’re on the steepest part of the curve' and that access will become an 'ever receding point.'
  39. 35 score Twitter/X · Quick skim · 1 min Joe Hudson, who coaches OpenAI's research team and advises Sam Altman and leaders… Joe Hudson, who coaches OpenAI's research team and advises Sam Altman and leaders at Apple and Google, is credited in Lenny Rachitsky's post (tweeted 2026-06-26) with identifying 'emotional clarity' as the key predictor of success in AI-forward teams.
  40. 35 score Twitter/X · Quick skim · 1 min On 2026-06-26, Jeremy Allaire (@jerallaire) announced Circle's support for… On 2026-06-26, Jeremy Allaire (@jerallaire) announced Circle's support for Proof's launch of x401, an open protocol to verify who authorized AI agents' actions in the so-called 'agentic economy'.
  41. 35 score ArXiv · Quick skim · 1 min Fast algorithms for learning a Gaussian under halfspace truncation with optimal sample complexity For any ε>0 and dimension d the authors give an efficient algorithm that learns a Gaussian truncated by an unknown halfspace using n = Õ(d^2/ε^2) samples and with runtime dominated by computing the empirical covariance matrix (paper claims this is optimal in d and ε).
  42. 35 score ArXiv · Quick skim · 1 min Autoregressive Boltzmann Generators ArBG introduces an autoregressive Boltzmann Generator framework that removes normalizing-flow topological constraints, enables sequential inference-time interventions, and scales via LLM-style architectures.
  43. 35 score ArXiv · Quick skim · 1 min All you need is log The paper proves that any functional of W-tuples of distributions that is data-processing monotone and additive on independent products equals a positive integral of multi-way coincidence divergences C_α(π_1,...,π_W) := -log ∫ π_1^{α_1}⋯π_W^{α_W} with ∑_k α_k = 1; the α-parameter space has four necessary strata: simplex interior, mixed-sign exponent cones, a tropical boundary (max-divergences), and pairwise KL edges at simplex vertices.
  44. 35 score Twitter/X · Quick skim · 1 min @tszzl claims the popular critique that an unofficial AI licensing regime is… @tszzl claims the popular critique that an unofficial AI licensing regime is 'slowing down innovation' ignores how quickly AI is moving; the Mythos incident may have accelerated oversight but such intervention was inevitable and 'earlier is better' amid exponential growth (posted 2026-06-26).
  45. 35 score Twitter/X · Quick skim · 1 min On 2026-06-26 Sarvesh Shrivastava (@bloggersarvesh) posted a playbook claiming he… On 2026-06-26 Sarvesh Shrivastava (@bloggersarvesh) posted a playbook claiming he could reach $100k/month in 90 days using Claude + SEO if he woke up bankrupt.
  46. 35 score ArXiv · Quick skim · 1 min BOWConnect (Raxit et al., accepted to IROS 2026, published 2026-06-25) integrates… BOWConnect (Raxit et al., accepted to IROS 2026, published 2026-06-25) integrates Bayesian Optimization over Windows (BOW) as a learned steering function inside a bidirectional parallel kinodynamic planner to learn local cost maps and guide constraint-aware control sampling.
  47. 35 score ArXiv · Quick skim · 1 min LA4VLA: Learning to Act without Seeing via Language-Action Pretraining LA4VLA constructs LA4-33K, a dataset of 33,000 Language-Action (LA) episodes by decomposing expert demonstration trajectories into atomic action segments paired with low-level action descriptions, created without additional robot data collection.
  48. 35 score ArXiv · Quick skim · 1 min RouterVLA: Turning Smoke Tests into Supervision for Heterogeneous VLA Selection RouterVLA uses outcome-disjoint cross-fitting on 34,752 LIBERO-Plus rollout records to build probe profiles for frozen VLA experts and raises held-out success from 0.4686 to 0.6149 (a +14.64 percentage-point gain) using a transparent probe-success rule.
  49. 35 score Twitter/X · Quick skim · 1 min On 2026-06-26 @cryptopunk7213 claims OpenAI's new GPT-5.6 'Sol' has officially… On 2026-06-26 @cryptopunk7213 claims OpenAI's new GPT-5.6 'Sol' has officially taken the #1 spot and "beats Mythos" on coding tasks.
  50. 35 score ArXiv · Quick skim · 1 min Data-Driven Duration Management -- Term Structure Forecasting Using Machine Learning Neural networks consistently beat classical term‑structure methods (Dynamic Nelson‑Siegel, PCA) on both U.S. Treasury and ECB zero‑coupon bond forecasts and in downstream portfolio performance; evaluation used RMSE, MAE, directional accuracy plus an economic bond‑trading metric (Lausser et al., 2026‑06‑25).
  51. 35 score ArXiv · Quick skim · 1 min Asymptotically Optimal Learning for Parametric Prophet Inequalities Characterized the optimal full-information asymptotic competitive ratio for i.i.d. rewards from an exponential-type parametric family (unknown θ): for unbounded-support distributions the limit equals ((θ/(θ - c_+))^{c_+/θ}) / Γ(1 - c_+/θ), while for bounded-support power-family the limit is 1.
  52. 35 score ArXiv · Quick skim · 1 min XMSE-Aware Adaptive Empirical Bayes Estimation Proposes an "XMSE-aware mixed estimator" (2026-06-25, Chen & Zheng) that linearly interpolates between maximum likelihood (ML) and a kernel-based empirical Bayes (EB) estimator; the fixed-weight excess mean squared error (XMSE) is a scalar quadratic, yielding a closed-form oracle mixing weight that is provably no worse than both ML and the base EB at the XMSE scale.
  53. 35 score ArXiv · Quick skim · 1 min Beyond Global Divergences: A Local-Mass Perspective on Bayesian Inference Introduces two tools: Mass Index (records polynomial and logarithmic decay scales of local mass) and regularised extended KL (RE-KL), a set-localised divergence that admits singular components.
  54. 35 score ArXiv · Quick skim · 1 min SAM2Matting: Generalized Image and Video Matting SAM2Matting (Ruiqi Shen, Guangquan Jie, Chang Liu, Henghui Ding; arXiv 2026-06-25; ECCV 2026 extended) is a tracker-to-matting framework that augments foundational VOS trackers (e.g., SAM2, SAM3) with a region-proposal bridge and dedicated matting heads, decoupling temporal tracking from fine-grained matting.
  55. 35 score ArXiv · Quick skim · 1 min RayPE: Ray-Space Positional Encoding for 3D-Aware Video Generation RayPE injects per-token 6D Plucker coordinates additively into queries and keys of self-attention (with a query/key flip) so the symmetric identity matches the Plucker reciprocal product; the resulting attention score cleanly decomposes into a content term, a geometry term, and two cross-terms, each found necessary by experiments.
  56. 35 score ArXiv · Quick skim · 1 min PhysiFormer: Learning to Simulate Mechanics in World Space PhysiFormer (Yiming Chen, Yushi Lan, Andrea Vedaldi; ArXiv 2026-06-25) is a diffusion transformer that samples future 3D mesh vertex trajectories in world coordinates from initial vertex positions, velocities, and material type (rigid or elastic) using a denoising diffusion process directly in coordinate space.
  57. 35 score ArXiv · Quick skim · 1 min DnA: Denoising Attention for Visual Tasks DnA (Denoising Attention) uses a positive query to select class-relevant image features and a negative query to select closely associated but irrelevant features, then projects their interactions into two distinct subspaces with larger principal angles to promote subspace separation and improved discriminability.
  58. 35 score ArXiv · Quick skim · 1 min World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays Recurrent Generative Replay (REGEN) leverages World Action Models (WAMs) to synthesize pseudo-replay trajectories by recursively querying a generative world+action model conditioned on prior task instructions and current-task observations; evaluated in both simulation and real-world robot manipulation (Govind et al., arXiv 2026-06-25).
  59. 35 score ArXiv · Quick skim · 1 min Bridging Talk and Thought: Understanding Dialogue Dynamics Across Collaborative Problem-Solving Contexts Zhengyuan Liu, Stella Xin Yin, Min-Yen Kan, and Nancy F. Chen (published 2026-06-25) introduce a hierarchical two-layer coding scheme that integrates cognitive and non-cognitive problem solving with explicit metacognitive regulatory mechanisms.
  60. 35 score ArXiv · Quick skim · 1 min LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank LLM-based pipeline achieved up to 91% document-level precision for eligibility decisions, operating conservatively to minimize false acceptance (paper reports high precision at the document level).
  61. 35 score ArXiv · Quick skim · 1 min Mapping Political-Elite Networks in Europe with a Multilingual Joint Entity-Relation Extraction Pipeline Modular open-weight multilingual pipeline that builds signed, temporal knowledge graphs from news using span-based NER, a three-stage linking cascade mapping mentions to language-independent Wikidata IDs, and an ontology-constrained mixture-of-experts model with guided decoding.
  62. 35 score ArXiv · Quick skim · 1 min Blackwell Approachability and Gradient Equilibrium are Equivalent Proves algorithmic equivalence between Gradient Equilibrium (GEQ) and Blackwell approachability: each can be solved using a black-box oracle for the other with no asymptotic loss in the oracle's error rate, and the reductions are efficient.
  63. 35 score ArXiv · Quick skim · 1 min DanceOPD: On-Policy Generative Field Distillation DanceOPD (2026-06-25) introduces on-policy generative field distillation for flow-matching models: each sample is routed to one capability-specific velocity field, the student queries a low-noise student-induced state, and training uses a simple velocity MSE objective to learn composition of capabilities.
  64. 35 score ArXiv · Quick skim · 1 min Simulation-based inference for rapid Bayesian parameter estimation in epidemiological models: a comparison with MCMC SBI (neural posterior estimation) calibrated a mechanistic SECIR COVID‑19 model on Germany 2020 ICU-occupancy data: for 31-day inference windows SBI matched MCMC posteriors while running ~60–70 seconds on a single GPU versus ~1,000 seconds for MCMC (CPU); for a 201-day reconstruction SBI averaged ~157 seconds vs >19,000 seconds for MCMC.
ArXiv · 1 min Signal

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

VISE (Visual Invariance Self-Evolution) is a purely unsupervised, single-model self-evolving LMM framework that targets 'visual under-conditioning' by directly regularizing visual conditioning with two invariance-based rewards: a geometric invariance reward (spatial consistency under known transformations) and a semantic invariance reward (penalizes evidence-agnostic generation when predicted regions are perturbed).

VISE introduces an unsupervised self-evolving method to fix visual under-conditioning in large multimodal models by enforcing visual attention through geometric invariance (consistency under spatial transforms) and semantic invariance (detecting absence of evidence when regions are perturbed). Trained on raw unlabeled images without specialist roles or external reward models, VISE substantially improves captioning (±CIDEr gains shown) and reduces hallucination across 18 benchmarks, evaluated with Qwen3-VL-2B and multiple model families (ECCV 2026).

On 18 benchmarks using Qwen3-VL-2B as the base model, VISE yields large gains: +16.85 CIDEr on COCO and +19.66 CIDEr on TextCaps, reduces object hallucination by 5.0 Chair-I points, and generalizes across four model families and scales; code and models are available at https://mbzuai-oryx.github.io/VISE (ECCV 2026).
Open reader
Worth reading

Useful context and follow-up reading when you have more time.

25 items
1 ArXiv 2026-06-25 1 min read
Open

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

Why it matters

VISE (Visual Invariance Self-Evolution) is a purely unsupervised, single-model self-evolving LMM framework that targets 'visual under-conditioning' by directly regularizing visual conditioning with two invariance-based rewards: a geometric invariance reward (spatial consistency under known transformations) and a semantic invariance reward (penalizes evidence-agnostic generation when predicted regions are perturbed).

  • On 18 benchmarks using Qwen3-VL-2B as the base model, VISE yields large gains: +16.85 CIDEr on COCO and +19.66 CIDEr on TextCaps, reduces object hallucination by 5.0 Chair-I points, and generalizes across four model families and scales; code and models are available at https://mbzuai-oryx.github.io/VISE (ECCV 2026).

VISE introduces an unsupervised self-evolving method to fix visual under-conditioning in large multimodal models by enforcing visual attention through geometric invariance (consistency under spatial transforms) and semantic invariance (detecting absence of evidence when regions are perturbed). Trained on raw unlabeled images without specialist roles or external reward models, VISE substantially improves captioning (±CIDEr gains shown) and reduces hallucination across 18 benchmarks, evaluated with Qwen3-VL-2B and multiple model families (ECCV 2026).

Authors: Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar...
2 ArXiv 2026-06-25 1 min read
Open

The Geometry of Updates: Fisher Alignment at Vocabulary Scale

Why it matters

Representation metrics (e.g., CKA) are non-identifiable for transfer in shared-output-head settings: models can share identical activations yet have orthogonal head updates. The paper proves an identity that head Fisher alignment equals the cosine between kernel mean embeddings in the joint activation–error space.

  • FisherSketch computes that cosine in a single streaming pass, making vocabulary-scale head Fisher alignment practical (K=128 or 256) with a 16 KB task signature (m=4096) and a 192 KB per-task streaming state. Experiments and diagnostics, including Llama-3.1-8B verbalizer-shift tests, show FisherSketch distinguishes tasks when activation similarity cannot; paper accepted to ICML 2026.

The paper addresses source selection for LLM families with shared vocabularies by characterizing update geometry without materializing Fisher matrices. It derives that head Fisher alignment is the cosine between kernel mean embeddings in the joint activation–error space, and introduces FisherSketch to estimate this in one streaming pass (16 KB signature, 192 KB state). Validations including Llama‑3.1‑8B show FisherSketch is informative where representation metrics fail.

Authors: John Sweeney
3 ArXiv 2026-06-25 1 min read
Open

LMs as Task-Specific Knowledge Bases: An Interpretability Analysis

Why it matters

Elhelo, Globerson, and Geva (ArXiv preprint, 2026-06-25) show language models encode factual knowledge in a task-specific way: behavioral experiments find facts acquired on one task frequently fail to co-emerge on others during training.

  • Mechanistic parameter-localization experiments reveal distinct parameter subsets support the same fact across tasks, and chain-of-thought prompts partly work by engaging task-specific parameters beyond those tied to the evaluation task—undermining the 'single source' knowledge-base analogy and raising reliability/controllability concerns.

Language models' factual knowledge is studied via behavioral and mechanistic analyses (Elhelo et al., arXiv 2026-06-25). The authors find facts learned on one task often do not appear on others, and parameter-localization shows different parameter subsets encode the same fact for different tasks. Chain-of-thought partly succeeds by activating task-specific parameters. Summary based on the abstract (full text not accessed).

Authors: Amit Elhelo, Amir Globerson, Mor Geva
4 Twitter/X 2026-06-26 2 min read
Open

Frontier models concentrate economic value

Why it matters

Frontier models concentrate economic value: customers currently pay for Anthropic and OpenAI while firms like Mistral and XAI are falling behind, producing winner-take-most dynamics and (in the most valuable markets) potentially winner-take-all outcomes.

  • The U.S. licensing regime will motivate China to develop frontier models because switching providers is often easy; the author estimates the U.S.–China R&D gap is under one year, so China could close it and capture customers if it offers the best models.
  • Attempts to ban or sanction Chinese models would likely fail in practice because digital borders are porous and a variety of resellers/wrappers can enable access; analogous EU attempts to regulate U.S. tech became ineffective and risk leaving regulators with outdated technology.
  • ‘Frontier’ is ambiguous: benchmark ranking depends on test-time compute (potentially requiring weeks or months), simple thresholds like FLOPs or parameter counts are imperfect, and with claims such as Andrej Karpathy expecting AGI in <1B-parameter models, the practical result could be licensing that eventually covers most models and model-capable hardware.

John Loeber (post 2026-06-26) argues that the new U.S. AI licensing regime will crystallize a contest over 'frontier' models because economic value concentrates at the frontier—customers pay for Anthropic and OpenAI while Mistral and XAI lag—creating winner-take-most (and sometimes winner-take-all) markets. He warns this will spur China to race to supply the best models, noting the R&D gap may be under one year and switching providers is often easy. Loeber says U.S. bans or sanctions on Chinese models would be ineffective given porous digital borders and resellers/wrappers, and that regulation without winning the technological race is a Red Queen's race. He also highlights that 'frontier' is hard to define—benchmarks depend on test‑time compute, FLOPs/parameter thresholds are imperfect, and claims like Karpathy’s (<1B parameters for AGI) imply licensing could ultimately extend to most models and hardware.

By @johnloeber
5 ArXiv 2026-06-25 1 min read
Open

Any system that outputs one member model's answer is upper-bounded by accuracy ≤…

Why it matters

Any system that outputs one member model's answer is upper-bounded by accuracy ≤ 1 − β, where β is the rate that every model is wrong on the same query; a Clopper–Pearson bound on β gives a finite-sample certificate on the maximum possible gain from routing, voting, or cascades.

  • Empirical evaluation across 67 models from 21 providers: on open-ended mathematics observed β = 0.052 versus β = 0.023 under a 67-model Gaussian copula (≈2.5× underpricing; 90% CI 1.7–3.4, k = 17); execution-graded code showed β = 0.079.
  • Re-asking GPQA-Diamond questions in free-response increased co-failure (β = 0.127) with a five-judge panel κ = 0.73–0.92; combining models rarely outperforms the single best model on checkable tasks without strong query-level routing, and gains come from models failing on different questions (heterogeneity), not just adding more models.

Chen shows a concrete "co-failure" ceiling for multi-model LLM systems: any policy that returns a single model's answer cannot exceed accuracy 1−β, where β is the rate all models fail on the same query. Using Clopper–Pearson finite-sample bounds and a 67-model empirical pool, the paper finds higher all-wrong tails than Gaussian-copula models predict (e.g., β=0.052 vs 0.023 on open math), and demonstrates that ensemble gains require heterogeneous failure patterns and query-level routing signals; full text not available, summary based on the abstract.

Authors: Josef Chen
6 ArXiv 2026-06-25 1 min read
Open

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs

Why it matters

RiVER (Ranking-induced VERifiable) trains LLMs using deterministic execution feedback as continuous rewards on score-based tasks, addressing 'scale dominance' and 'frequency dominance' via instance-wise calibrated reward shaping that emphasizes top-ranked solvers while keeping bounded feedback for other valid solutions; training used 12 AtCoder Heuristic Contest tasks.

  • On benchmark ratings, RiVER improves ALE rating rank for Qwen3-8B by 8.9% and for GLM-Z1-9B-0414 by 9.4% compared to baselines.
  • RiVER—trained exclusively on score-based (no ground-truth) tasks—also transfers to exact-solution benchmarks, boosting backbones on LiveCodeBench and USACO by absolute average improvements of 2.4% and 3.5%; by contrast, models trained on raw execution scores improved ALE but failed to transfer.

RiVER (Ranking-induced VERifiable) uses deterministic execution scores as continuous supervision and a calibrated, instance-wise reward shaping that emphasizes top-ranked candidates to avoid scale and frequency dominance. Trained on 12 AtCoder Heuristic Contest tasks, RiVER raises ALE rating ranks for Qwen3-8B and GLM-Z1-9B-0414 by 8.9% and 9.4% and yields 2.4%/3.5% absolute gains on LiveCodeBench/USACO, despite no ground-truth solutions.

Authors: Yingyu Lin, Qiyue Gao, Nikki Lijing Kuang...
7 ArXiv 2026-06-25 1 min read
Open

Hallucination in World Models is Predictable and Preventable

Why it matters

MMBench2: a 427-hour, 210-task visual world-modeling dataset with ground-truth actions, rewards, and live simulators; authors train a 350M-parameter generative world model on it.

  • Three distinct hallucination modes are identified—perceptual, action-marginalized, and scene-diverging—and the paper develops three lightweight data-centric signals that accurately predict where the model will fail.
  • Mitigation via coverage-aware sampling and using the hallucination predictors as curiosity rewards enables data-efficient finetuning: the pretrained world model adapts to entirely unseen environments with as few as 50 real environment trajectories, supporting the claim that hallucination is primarily a data-coverage issue.

Hansen and Wang introduce MMBench2 (427 hours, 210 tasks) and train a 350M-parameter world model to study hallucination in generative, action-controllable rollouts. They categorize failures into perceptual, action-marginalized, and scene-diverging modes, develop three predictive signals, and use coverage-aware sampling plus curiosity-driven data collection to finetune on new environments with as few as 50 trajectories, arguing hallucination arises from data-coverage gaps.

Authors: Nicklas Hansen, Xiaolong Wang
8 ArXiv 2026-06-25 1 min read
Open

Multilingual Reasoning Cascades Need More Context

Why it matters

A simple, training-free context-aware cascade—giving the final translation module the original question, the English-translated question, and the English reasoning trace—yields strong gains for open-ended generation across nine multilingual benchmarks, three backbone models, and 285 high-/mid-/low-resource languages (ArXiv, 2026-06-25).

  • Ablations show the original-language question supplies most of the helpful context; preserving the original user question until the end of the pipeline is the primary actionable strategy to mitigate error propagation.
  • The paper diagnoses translation cascades as structurally lossy—stages discard cues for cultural grounding, register, and disambiguation—motivating redesign of information flow in multilingual reasoning systems.

The authors introduce a context-aware translation cascade that appends the original question, its English translation, and the English reasoning trace to the final translation step. Evaluated on nine multilingual benchmarks with three backbone models across 285 languages, the training-free intervention markedly improves open-ended generation. Ablations attribute most gains to preserving the original-language question and argue for rethinking cascade information flow.

Authors: Arnav Mazumder, Dengjia Zhang, Shuyue Stella Li...
9 ArXiv 2026-06-25 1 min read
Open

Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance

Why it matters

Proposes a training-free, plug-and-play feature self-guidance that disperses internal features during batch generation and uses a manifold regularization step to project dispersed features back onto the data manifold, aiming to mitigate diversity collapse in pretrained flow models.

  • Demonstrates significant improvements in sample diversity while preserving fidelity across conditional flow tasks (multi-step and few-step text-to-image, depth-to-image, and reference-image generation) with only marginal inference cost and no external reward models.
  • Paper by Pradhaan S Bhat, Rishubh Parihar, Abhijnya Bhat, and R. Venkatesh Babu; accepted to ECCV 2026, posted on arXiv 2026-06-25 as arXiv:2606.27371v1 (project page: https://dont-settle-at-the-mode.github.io/).

A training-free, plug-and-play feature self-guidance method mitigates diversity collapse in pretrained flow models by dispersing internal features during batch generation and then projecting them back to the data manifold via manifold regularization. The approach avoids external reward models, adds marginal inference overhead, and reportedly increases generation diversity while maintaining fidelity across multiple conditional flow tasks.

Authors: Pradhaan S Bhat, Rishubh Parihar, Abhijnya Bhat...
10 ArXiv 2026-06-25 1 min read
Open

VibeAct: Vibration to Actions for Contact-Rich Reactive Robot Dexterity

Why it matters

VibeAct embeds piezoelectric microphones in a dexterous robot hand, records vibro-acoustic data via teleoperation, and replays recordings in a calibrated digital clone to auto-label per-finger contact and slip for training a tactile estimator.

  • Trained policies use the estimator's contact and continuous slip-magnitude channels (not raw audio) and, across five contact-rich tasks (regrasping, in-hand reorientation, insertion), VibeAct outperforms a proprioception+point-cloud baseline in simulation and transfers to a real hand-arm platform, improving success rates.

VibeAct tackles occluded, fast contact events by using high-bandwidth piezoelectric microphones on a dexterous hand, teleoperating to collect vibro-acoustic traces, and replaying them in a calibrated digital clone to automatically label per-finger contact and slip. A tactile estimator maps real waveforms to contact/slip; RL policies are trained in simulation on those labels (avoiding raw-audio sim). Evaluated on five contact-rich tasks, VibeAct yields the largest gains on sustained reactive control where the continuous slip-magnitude channel is most informative, and the policies transfer to physical hardware. (Authors: Yuemin Mao et al., arXiv 2026; project: https://vibeact.github.io/.)

Authors: Yuemin Mao, Uksang Yoo, Jean Oh...
11 ArXiv 2026-06-25 1 min read
Open

How Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple Mitigation

Why it matters

The paper introduces a four-dimension diagnostic (tokenization cost, predictive uncertainty/surprisal, semantic robustness, context sensitivity) and evaluates it on three corpora: 17th‑century Italian (1610–1689), 19th‑century I Promessi Sposi (high-exposure control), and 18th‑century Russian civil print books.

  • Tokenization inflation is comparable for Russian and early modern Italian (25–30%), but 17th‑century Italian is on average 2.4× more surprising than modern Italian (academic prose up to 3.2×); Russian shows only a modest surprisal increase.
  • Semantic representations remain robust (embedding similarity >0.85 across datasets), and a minimal temporal context prompt reduces historical surprisal by ~60%, offering a simple, model-agnostic mitigation for generative tasks.

The paper evaluates LLMs' handling of historical language by proposing a four-part diagnostic (tokenization cost, predictive surprisal, semantic robustness, context sensitivity) and testing on 17th‑century Italian (1610–1689), 19th‑century I Promessi Sposi (control), and 18th‑century Russian. It finds 25–30% tokenization inflation, 17th‑century Italian 2.4× surprisal (3.2× in academic prose), embeddings >0.85, and a temporal prompt cutting surprisal ~60%.

Authors: Maria Levchenko
12 ArXiv 2026-06-25 1 min read
Open

Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders

Why it matters

Top-k SAE augmented with two sparsity regularizers—an ℓ1 penalty on off-support (unselected) units and a scale-invariant ℓ1/ℓ2 ratio on batch-active units—by Jacquier, Vakalopoulou, and Hosseini (published 2026-06-25) increases monosemanticity across two datasets, three vision foundation models, and a range of k while preserving reconstruction quality.

  • The ℓ1/ℓ2 ratio concentrates information into fewer effective latents, making reconstruction more robust to the inference-time choice of k and improving small-budget linear probing; both penalties act on activations before Top-k and are applied only to units selected at least once within the batch.

Top-k sparse autoencoders, which fix a hard budget k, are augmented with two compatible sparsity regularizers: an ℓ1 penalty on off-support units and a scale-invariant ℓ1/ℓ2 ratio on batch-active units, both applied before Top-k. Evaluated on two datasets and three vision foundation models, these regularizers consistently improve monosemanticity without harming reconstruction, and the ℓ1/ℓ2 variant further concentrates code and improves robustness to inference-time k and small-budget linear probing. Summary based on the paper abstract.

Authors: Nathanaël Jacquier, Maria Vakalopoulou, Mahdi S. Hosseini
13 ArXiv 2026-06-25 1 min read
Open

Decision-Aligned Evaluation of Uncertainty Quantification

Why it matters

Authors Annika Schneider, Tommy Rochussen, Joshua Stiller, and Vincent Fortuin (ArXiv:2606.26990v1, published 2026-06-25) introduce a decision-alignment criterion and show that common UQ metrics—negative log-likelihood (NLL) and expected calibration error (ECE)—can be misaligned with downstream decision utility.

  • They propose 'prior-weighted utility metrics', a special class of proper scoring rules that incorporate downstream priors to produce decision-aligned uncertainty evaluation.
  • Across benchmark experiments and real-world case studies reported in the paper, the proposed metrics consistently align with realized decision utility, whereas conventional metrics often fail or encode pathological prior beliefs about the task.

Decision-alignment introduces a criterion to judge whether uncertainty quantification (UQ) evaluation metrics (e.g., NLL, ECE) meaningfully reflect downstream decision utilities. The authors diagnose misalignment and pathological implicit priors in many standard metrics, and propose prior-weighted utility metrics—a family of proper scoring rules. Benchmarks and real-world case studies (ArXiv:2606.26990v1, 2026-06-25) show these metrics better predict realized decision utility.

Authors: Annika Schneider, Tommy Rochussen, Joshua Stiller...
14 ArXiv 2026-06-25 1 min read
Open

Scalable Behavior Cloning with Open Data, Training, and Evaluation

Why it matters

ABC-130K: released the largest open-source teleoperation dataset to date with 3,500 hours across ~130K episodes covering 195 manipulation tasks (paper posted 2026-06-25).

  • Open-source stack: published hardware designs, training infrastructure, a simulation pipeline, plus 400 hours of sim-teleop data and a co-training recipe that yields correlated simulation↔real evaluations for low-cost ablation.
  • Empirical study: compared Diffusion Transformers (DiT) and Vision-Language-Action (VLA) models, validated on real-world dexterous tasks (e.g., box folding and extracting credit cards from wallets) and released a reproducible toolkit.

ABC presents a fully open-source behavior-cloning stack centered on ABC-130K (3,500 hours, ~130K episodes, 195 tasks), plus hardware, training, and simulation releases. They add 400 hours of sim-teleop data and a co-training recipe that correlates sim and real evaluations, compare DiT and VLA architectures, and show policies succeeding on dexterous manipulation, enabling reproducible research.

Authors: Arthur Allshire, Himanshu Gaurav Singh, Ritvik Singh...
15 ArXiv 2026-06-25 1 min read
Open

When are likely answers right? On Sequence Probability and Correctness in LLMs

Why it matters

Across prompt–answer pairs within a fixed dataset, higher sequence probability often predicts correctness: Zenn & Geiping (2026) quantify this alignment at four levels (decoding methods, hyperparameters, prompt–answer pairs, repeated responses) in a 38‑page preprint (10 pages main, 28 pages appendix).

  • Changing decoding method or hyperparameters to increase sequence probability does not reliably improve accuracy, and sequence probability is a poor indicator of correctness for repeated responses — with direct implications for decoding choices, self‑consistency, and verifier‑free self‑improvement.

The paper studies when sequence probability (the model's conditional probability of a continuation) aligns with correctness by measuring this relationship across decoding methods, hyperparameter settings, prompt–answer pairs, and repeated responses. Key findings: within a fixed dataset higher sequence probability often correlates with correctness, but raising sequence probability via decoding changes usually does not boost accuracy, and sequence probability fails to predict correctness for repeated responses. The results clarify limits of probability‑based decoding and guide self‑consistency and verifier‑free improvement strategies; full preprint (PDF) is available.

Authors: Johannes Zenn, Jonas Geiping
16 ArXiv 2026-06-25 1 min read
Open

Error-Conditioned Neural Solvers

Why it matters

Error-Conditioned Neural Solvers (ENS) (Jiang et al., arXiv 2026-06-25) feed the PDE residual field into the network at every iteration so the model learns an update policy to correct its own errors, avoiding expensive classical hybrid optimizers (gradient descent / Gauss–Newton) and their instability.

  • Across four PDE families, ENS attains the highest prediction accuracy in the majority of settings, with improvements up to 10× on turbulent Kolmogorov flow, and it generalizes under distribution shift (zero-shot parameter changes and cross-equation transfer), especially in ill-conditioned regimes where residual minimization is unreliable.

Error-Conditioned Neural Solvers (ENS) address limitations of neural surrogates and hybrid residual-minimization methods by passing the PDE residual field as an explicit input at each iteration so the network learns a correction policy instead of numerically minimizing the residual. The authors prove residual minimization can be a poor proxy in ill-conditioned systems and empirically show ENS outperforms hybrids across four PDE families—up to 10× on Kolmogorov flow—while reducing compute and improving robustness to distribution shift.

Authors: Haina Jiang, Liam Wang, Peng-Chen Chen...
17 ArXiv 2026-06-25 1 min read
Open

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

Why it matters

Proposes a self-evolving training framework (published 2026-06-25) that decomposes a unified LMM into three internal roles—Proposer (generates visual questions), Solver (answers and evaluates), and Generator (synthesizes images)—and trains using only unlabeled images with no human annotations, preference labels, or external reward/judge models.

  • Introduces Solver Token Entropy (STE), a continuous token-level difficulty signal to stabilize learning when sample-level consistency is unreliable, and a multi-scale internal evaluation for generation combining question–answer fidelity with cycle-consistent captioning to couple understanding and generation.
  • Reports consistent gains across eight understanding metrics and across BLIP3o, rectified-flow BAGEL, and autoregressive VARGPT-v1.1; on BAGEL achieves +3.5% absolute on MMMU and improves GenEval image-generation score from 82% to 85%. Code and models are publicly released.

The paper presents Ask–Solve–Generate, a self-evolving framework that lets unified large multimodal models improve vision understanding and image generation using only unlabeled images. A Proposer, Solver and Generator interact via self-consistency rewards; Solver Token Entropy (STE) stabilizes training, and a multi-scale QA+cycle-caption evaluator assesses generations. The approach yields consistent improvements (e.g., BAGEL +3.5% MMMU, GenEval 82%→85%) and generalizes across BLIP3o, BAGEL, and VARGPT-v1.1. Full text and code were released on arXiv.

Authors: Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar...
18 ArXiv 2026-06-25 1 min read
Open

Bridging Performance and Generalization in Reinforcement Learning for Agile Flight

Why it matters

The authors combine task-aware switching (based on learning progress) with a physically informed procedural track generator to produce a fast, robust zero-shot generalist policy for drone racing, achieving a 7.4x improvement in generalization over prior state-of-the-art while keeping competitive racing speeds and requiring no test-time adaptation.

  • Results are validated in both simulation and real-world experiments, including a challenging vision-based end-to-end control setting that operates without explicit state estimation—an operating regime where all prior approaches reportedly fail to generalize.

Autonomous drone racing faces time-optimal control under actuation saturation and poor zero-shot generalization; the authors address this by combining task-aware switching (learning-progress-based) with a physically informed procedural track generator. The method yields a 7.4x generalization improvement versus prior work, preserves competitive speeds, and is validated in simulation and real-world, vision-only end-to-end tests. Summary based on the paper abstract.

Authors: Jonathan Green, Jiaxu Xing, Nico Messikommer...
19 Twitter/X 2026-06-26 1 min read
Open

cryptopunk7213 claims every AI model will fall into two groups — 'Brain' (highly…

Why it matters

cryptopunk7213 claims every AI model will fall into two groups — 'Brain' (highly capable) and 'Executor' (workers) — and whoever builds the smart‑router platform on top will capture outsized value; 'Brain' models will cost ~5–10× more per token and Fortune 100 companies will pay for them.

  • About 80% of token consumption will go to execution/worker models, so the winner for execution will be determined by cost‑per‑token; Chinese open‑source models (Qwen, DeepSeek, MiniMax, GLM, Kimi) are positioned as current cost leaders because they can run locally or via cloud catalogs.
  • UBS (reported via Rohan Paul) finds 60% of companies are moving to cheaper models or open source; enterprises face extreme bills (users up to $35K/month, teams exceeding quotas by 200%) and are using model routing to send easy tasks to cheaper models while reserving premium models for hard reasoning, code, and long‑context work.

cryptopunk7213 argues AI models will split into 'Brain' and 'Executor' tiers and that smart‑router platforms will capture the most value. 'Brain' models will run at ~5–10× token cost for Fortune 100 use, while ~80% of tokens go to cheap executor models — often Chinese open‑source options (Qwen, DeepSeek, MiniMax, GLM, Kimi).

By @cryptopunk7213
20 ArXiv 2026-06-25 1 min read
Open

Ribbon: Scalable Approximation and Robust Uncertainty Quantification

Why it matters

Ribbon (Gibson, Tipton, Rumsey, Klein; arXiv 2026-06-25) approximates Dirichlet-reweighted (Bayesian/weighted-likelihood) bootstrap by using an influence-function linearization around a single fitted model, replacing repeated refitting with post-hoc linear algebra.

  • Ribbon provides a calibrated Dirichlet-reweighting family with a general concentration parameter that can be tuned on validation data; it is asymptotically equivalent to a flat-prior Laplace approximation under correct likelihood specification and recovers the robust sandwich covariance under misspecification.
  • On synthetic regression, MNIST classification, and California Housing benchmarks Ribbon delivered competitive predictive performance and improved calibration in several settings while avoiding repeated model retraining.

Ribbon (Gibson et al., arXiv 2026-06-25) targets scalable predictive uncertainty for complex or misspecified models by approximating the Dirichlet-reweighted bootstrap (Bayesian/weighted-likelihood bootstrap) via an influence-function linearization around a single fitted model, avoiding costly posterior sampling or repeated refits. It supports a tunable concentration parameter, matches Laplace asymptotics under correct likelihoods, recovers sandwich covariances under misspecification, and (per the abstract) improves calibration on benchmarks; full text was not provided here.

Authors: Graham Gibson, John Tipton, Kellin Rumsey...
21 ArXiv 2026-06-25 1 min read
Open

Hamid Reza Firoozfar et al.

Why it matters

Hamid Reza Firoozfar et al. (published 2026-06-25; arXiv:2606.27314v1) introduce a comprehensive, mechanism-oriented taxonomy of indirect linguistic expressions (ILE) that categorizes the underlying encoding and recovery operations rather than communicative goals.

  • They embed the taxonomy into LLM prompts and evaluate it against four existing taxonomies plus a no-taxonomy baseline on 2,000 manually annotated TikTok and Bluesky posts using three LLMs, achieving the strongest document- and span-level detection performance.
  • The taxonomy yields an absolute improvement of 4.7% accuracy and 5.4% F1 over the best-performing benchmark; the work was submitted for ARR review for EMNLP 2026 and the abstract warns of profane/offensive content.

A mechanism-oriented taxonomy for indirect linguistic expressions (ILE) categorizes the operations that encode and recover meaning (e.g., algospeak, euphemisms, adversarial obfuscation) rather than speakers' intents. The authors integrate this taxonomy into LLM prompts and, on 2,000 annotated TikTok and Bluesky posts evaluated with three LLMs, report the best document- and span-level results—improving accuracy by 4.7% and F1 by 5.4% over prior taxonomies. Summary based on the abstract (full text not reviewed).

Authors: Hamid Reza Firoozfar, Mohammadsadegh Abolhasani, Reza Mousavi...
22 Twitter/X 2026-06-26 1 min read
Open

GPT‑5.6 will run on Cerebras hardware at up to 750 tokens per second; OpenAI says…

Why it matters

GPT‑5.6 will run on Cerebras hardware at up to 750 tokens per second; OpenAI says GPT‑5.6 Sol on Cerebras will be available in July 2026 with initial access limited to select customers.

  • OpenAI priced GPT‑5.6 per 1M tokens: Sol $5 input / $30 output; Terra $2.50 input / $15 output; Luna $1 input / $6 output; new caching rules include explicit cache breakpoints, a 30‑minute minimum cache life, cache writes billed at 1.25× the uncached input rate, and cache reads receiving a 90% cached‑input discount.
  • The author claims OpenAI’s strategy to own the AI hardware stack (Cerebras now, and a custom 'jalapeño' chip expected end of year 2026) will make intelligence much cheaper and faster and free compute to 'build a better model.'

GPT‑5.6 will run on Cerebras hardware at up to 750 tokens/sec (OpenAI plans Sol on Cerebras in July 2026) while introducing per‑1M‑token pricing for Sol/Terra/Luna and tightened prompt‑caching rules. The poster argues OpenAI owning the hardware stack—and Cerebras' upcoming 'jalapeño' chip later in 2026—will drive much faster, cheaper inference and accelerate model development.

By @cryptopunk7213
23 ArXiv 2026-06-25 1 min read
Open

Continual Robot Policy Learning via Variational Neural Dynamics

Why it matters

Proposes a continual learning framework that combines an analytical physics prior with a neural residual and a recurrent encoder that infers hidden, recurring dynamics from recent state-action trajectories; the inferred latent conditions condition both the residual dynamics model and the policy, and policy learning is done via differentiable simulation over sampled latent dynamics.

  • Real quadrotor experiments under changing wind: the policy recovers from recurring disturbances in ≈1 s (≈5× faster than online residual re-fitting) and reduces large-disturbance hover and tracking errors by 65.7% and 53.3%, respectively, versus state-of-the-art online adaptation.

A continual learning framework for robot control learns condition-aware dynamics by combining an analytical physics prior with a neural residual and a recurrent encoder that infers hidden, recurring dynamics from recent experience. Policies are trained in differentiable simulation over sampled latent conditions and at deployment use online inference to recover recurring disturbances. In quadrotor tests under changing wind, the approach achieves ~1 s recovery and major error reductions compared to SOTA.

Authors: Jiaxu Xing, Zhiyuan Zhu, Yunfan Ren...
24 ArXiv 2026-06-25 1 min read
Open

OctoSense: Self-Supervised Learning for Multimodal Robot Perception

Why it matters

OctoSense releases a 59-hour time-synchronized driving dataset (published 2026-06-25) containing stereo RGB + event cameras, LiDAR, thermal camera, IMU, RTK GPS, and proprioception (CAN bus and quadruped joint angles) across varied environments and times of day, including highly degraded-sensor scenarios.

  • The proposed model is a late-fusion masked autoencoder with modality-specific tokenizers and cached modality tokens at inference; representation computation is fast: 6.68 ms on an NVIDIA 5090 and 112 ms on an Orin NX.
  • Their multimodal self-supervised approach outperforms existing image-only foundation models on downstream tasks (optical flow, depth, semantic segmentation, ego-motion: translation, rotation, steering angle) and yields robust predictions at nighttime or under degraded sensing; dataset and code are open-source (project page link provided).

OctoSense presents an open-source multimodal sensor platform and a 59-hour time-synchronized driving dataset (stereo RGB, event cams, LiDAR, thermal, IMU, RTK GPS, proprioception). They train a late-fusion masked autoencoder with modality-specific tokenizers and cached tokens to handle heterogeneous frequencies/latencies. The model is real-time (6.68 ms / 112 ms) and outperforms image-only baselines on flow, depth, segmentation, and ego-motion, remaining robust in nighttime/degraded conditions. Summary based on the paper abstract (PDF link provided).

Authors: Anthony Bisulco, Jeremy Wang, Kostas Daniilidis...
25 Twitter/X 2026-06-26 1 min read
Open

Ornith-1.0 is an open-source family of agentic-coding LLMs with model sizes 9B…

Why it matters

Ornith-1.0 is an open-source family of agentic-coding LLMs with model sizes 9B Dense, 31B Dense, 35B MoE, and 397B MoE; all models are released under the MIT license and were post-trained on top of gemma4 and qwen3.5.

  • Reported benchmark scores include Terminal-Bench 2.1 = 77.5; SWE-Bench = 82.4 (verified), 62.2 (pro), 78.9 (multilingual); NL2Repo = 48.2; SWE Atlas = 41.2 (QnA), 42.6 (RF), 39.1 (TW); ClawEval = 77.1.
  • Ornith claims a novel RL-based 'self-improving' training that jointly optimizes task-specific scaffolds and solution rollouts, and the authors/observers suggest the 397B MoE may match or outperform Claude Opus 4.8.

Ornith-1.0 is an open-source family of agentic-coding LLMs (9B–397B, including MoE variants) post-trained on gemma4 and qwen3.5 that reports state‑of‑the‑art open-source coding scores (e.g., Terminal‑Bench 77.5, SWE‑Bench 82.4 verified). It uses an RL-based 'self‑improving' scaffold+rollout training technique and is released under the MIT license; authors claim parity or superiority to Claude Opus 4.8.

By @kimmonismus
Quick skim

Scan these for facts, links, or weak signals worth tracking.

39 items · open
1 Twitter/X 2026-06-26 1 min read
Open

On 2026-06-26, @johnloeber argues that in rivalries like the US vs China in AI…

Why it matters

On 2026-06-26, @johnloeber argues that in rivalries like the US vs China in AI, when one side slows the other 'goes even faster'—pushing harder to overtake or to establish a permanent lead.

  • He warns this dynamic could kickstart a Chinese state-capacity effort (a government-led push) to win the international AI market.

John Loeber (2026-06-26) argues that in competitive rivalries—exemplified by the US and China in AI—a slowdown by one side typically prompts the other to accelerate rather than relax. He predicts the gap could be treated as an opening, potentially triggering a Chinese state-capacity campaign to capture the international AI market.

By @johnloeber
2 ArXiv 2026-06-25 1 min read
Open

Language-Based Digital Twins for Elderly Cognitive Assistance

Why it matters

Proposes a language-based digital twin framework that leverages large language models plus stylometric cues and contextual metadata, and introduces a multi-head conditional variational autoencoder (cVAE) to jointly measure reconstruction quality and predict cognitive (MoCA) scores.

  • On the I-CONECT dataset, the digital twin preserves identity-specific conversational characteristics, achieves reconstruction and MoCA prediction errors comparable to real data, and outperforms baseline GPT-generated responses; paper posted on arXiv 2026-06-25 and accepted at PETRA 2026.

Language-based digital twins for elderly cognitive assistance propose a framework that leverages large language models, stylometric cues, and contextual metadata to mimic older adults' conversational behavior. The authors introduce a multi-head conditional variational autoencoder to measure reconstruction fidelity and predict MoCA scores. On the I-CONECT dataset the twins preserve identity-specific features, match real-data reconstruction and MoCA errors, and outperform GPT baselines.

Authors: Mohammad Mehdi Hosseini, Mohammad H. Mahoor, Hiroko H. Dodge
3 Twitter/X 2026-06-23 1 min read
Open

Mitchell Hashimoto announced on 2026-06-23 that his family is donating another…

Why it matters

Mitchell Hashimoto announced on 2026-06-23 that his family is donating another $400,000 to the Zig Software Foundation.

  • Hashimoto said he uses AI every day but highlighted that Zig has “one of the strongest anti-AI policies in open source,” adding “we disagree on some things, but respect doesn’t require agreement.”
  • @kiwicopple praised this stance and called that attitude “an increasingly rare mindset.”

Mitchell Hashimoto announced on 2026-06-23 that his family is donating another $400,000 to the Zig Software Foundation, calling Zig "exceptional software." He says he uses AI daily but supports Zig's "one of the strongest anti-AI policies in open source" and stresses "respect doesn’t require agreement." @kiwicopple called that an "increasingly rare mindset."

By @kiwicopple
4 Twitter/X 2026-06-26 1 min read
Open

@joshrauh asserts Gavin Newsom's premise that the wealthy pay lower tax rates…

Why it matters

@joshrauh asserts Gavin Newsom's premise that the wealthy pay lower tax rates than the rest 'is a lie,' citing Auten and Splinter showing the average tax rate rises from 2% for the lowest quintile to 45% for the top 0.01% of taxpayers.

  • On 2026-06-26 Gavin Newsom tweeted advocating a national 'billionaires tax' and a new social contract, claiming 10% of Americans own two-thirds of the wealth, wages have stagnated, the cost of living has skyrocketed, and the federal tax, corporate, and inheritance codes need an 'economic reset.'

@joshrauh criticizes Gavin Newsom's June 26, 2026 call for a national billionaires tax, saying Newsom's premise that the wealthy pay lower tax rates 'is a lie' and citing Auten & Splinter data (average tax rate 2% for the lowest quintile vs. 45% for the top 0.01%). Newsom argues 10% own two-thirds of wealth and calls for an economic reset.

By @joshrauh
5 Twitter/X 2026-06-26 1 min read
Open

swyx says they’ve been “scaling without slop” by working with aligned domain…

Why it matters

swyx says they’ve been “scaling without slop” by working with aligned domain experts to add coverage and co-organized their first AI Forward Deployed Engineering (FDE) miniconference with Basil Chatha at AI Engineer World’s Fair.

  • swyx claims OpenAI and Anthropic are launching “multi‑billion dollar services arms,” and argues FDE is one of the most in‑demand disciplines for Enterprise AI—though he personally has never done FDE.
  • AI Engineer World’s Fair 2026 runs June 29–July 2 in San Francisco (29 tracks, 300 speakers, 100 expo partners, 6,000+ attendees); the Forward Deployed track by Basil Chatha is on June 30 focusing on deploying agents in production.

swyx reports they’re “scaling without slop” by partnering with aligned domain experts and co-hosted an AI Forward Deployed Engineering miniconference with Basil Chatha at AI Engineer World’s Fair. He asserts OpenAI and Anthropic are launching multi‑billion‑dollar services arms, calls FDE essential for Enterprise AI adoption, and notes he has not personally worked as an FDE. The conference runs June 29–July 2, 2026 in San Francisco, with the FDE track on June 30.

By @swyx
6 Twitter/X 2026-06-25 1 min read
Open

On June 25, 2026 Governor Gavin Newsom announced California’s…

Why it matters

On June 25, 2026 Governor Gavin Newsom announced California’s “first-in-the-nation” dashboard to proactively track AI-related job-loss trends; the tool was developed with the University of California and the California Policy Lab and released under his recent executive order (GovPressOffice tweet).

  • The dashboard is described as an early-warning system to monitor, track, and anticipate job loss and is paired with a comprehensive data analysis intended to help prepare workers, small businesses, and communities for AI-driven disruption.
  • Author @krassenstein praised the launch as “the type of leadership we need,” warning that artificial intelligence will “completely upend the jobs market” and urging immediate forward-thinking action.

Governor Gavin Newsom announced on June 25, 2026 the launch of California’s first-in-the-nation dashboard to proactively track AI-related job-loss trends. Developed with the University of California and the California Policy Lab under his executive order, the dashboard is an early-warning system paired with comprehensive analysis to help prepare workers, small businesses, and communities.

By @krassenstein
7 Twitter/X 2026-06-26 1 min read
Open

At age 20, Turan Selvi says he was living with his parents in a small village in…

Why it matters

At age 20, Turan Selvi says he was living with his parents in a small village in Germany, peaked at 310 pounds, was depressed and barely sleeping.

  • He booked a flight to Los Angeles for 3 months and credits being around 'better people' and better energy for working harder and losing 'a ton of weight.'
  • He asserts that 'no one is coming to save you' and that past hardships forced him to obsess over independence and control of his life.

Turan Selvi recounts being 20, living with his parents in a German village at 310 pounds, depressed and sleepless. He says a three-month move to Los Angeles exposed him to better people and energy, spurred greater effort and major weight loss, and convinced him that hardship taught him self-reliance and control.

By @fromzerotomill
8 Twitter/X 2026-06-26 3 min read
Open

Tony's 'human in the loop = 20x pricing' play

Why it matters

Tony's 'human in the loop = 20x pricing' play: a buddy's end‑to‑end AI A/B‑testing and funnel product got no traction as $500/mo software but, repackaged identically as a human‑accountable service, started closing $15,000/month contracts.

  • Price anchoring drove rapid growth: a beauty‑tech founder launched the same product labeled at 2x the price ($2,000 vs $1,000), which tripled his business because buyers chose the $2K premium; Grey Goose used a similar physical/price anchor to move from bottom‑shelf to premium.
  • Zero switching costs between AI models: Jacky routinely exports personal context from Claude, uses Claude Code Max ($220/mo) as his daily driver, runs Kimi Code alongside, and leverages GLM (z.ai) on the free tier for large overnight UX/UI audits across every page and div.
  • Mind‑share concentration tactic: only ~3 winners dominate a category—dominate a tiny vertical early (example: Vessi owning disc golf shoes) so a small TAM is acceptable if you become #1 before the market scales.

Tony's thread lays out five actionable business plays and cautionary anecdotes from the week: repackaging identical AI software as a human‑backed service can jump price from $500/mo to $15K/month and close enterprise deals; deliberate price anchoring (e.g., $2K vs $1K) can 3x revenue by shifting buyers to a premium SKU; AI loyalty is collapsing—users move data between models freely, with stacks like Claude Code Max ($220/mo), Kimi, and free GLM doing heavy lifting such as overnight UX audits; category mindshare centralizes around roughly three winners, so owning a tiny vertical (Vessi in disc golf shoes) is a valid growth path; and both finance and brand work badly when incentives are misaligned—most hedge funds underperform the S&P over 10–15 years, and expensive brand redesigns can destroy SEO, forcing operational resets.

By @indexsy
9 Twitter/X 2026-06-26 2 min read
Open

Gina Acosta (@ginacostag_) published a "Top 20 Free AI Tools You Need in 2026"…

Why it matters

Gina Acosta (@ginacostag_) published a "Top 20 Free AI Tools You Need in 2026" list on 2026-06-26, categorizing tools across Productivity, Writing, Audio, App/Web dev, Coding, and Image/Design and naming examples like GetStudyPal, Perplexity, Google Gemini, ChatGPT, Grammarly, ElevenLabs, Replit, GitHub Copilot, MidJourney, Canva Magic Studio, and Ideogram.

  • She reports her AI agent built a fully working jobs board from scratch using Zero (zero.xyz), which she says provides access to ~14,000 real-world services so the agent could call live endpoints with no API keys, no signup, deliver real, filterable job listings on every load, and require no external dashboard or subscriptions (promo: $5 free credit).

Gina Acosta's 2026-06-26 post shares a Top 20 free AI tools roundup across productivity, writing, audio, development, coding, and design (examples: GetStudyPal, Perplexity, ChatGPT, ElevenLabs, Replit, Copilot, MidJourney). She also claims an AI agent built a live, filterable jobs board end-to-end by calling services through Zero (zero.xyz), which she says exposes ~14,000 services with no API keys or signup and includes a $5 credit promo.

By @ginacostag_
10 Twitter/X 2026-06-26 1 min read
Open

In 1999 Pontevedra's mayor removed street parking, eliminated most…

Why it matters

In 1999 Pontevedra's mayor removed street parking, eliminated most through‑traffic, and restricted car access to residents, deliveries, taxis and emergencies.

  • Store owners initially protested, but the changes made the city calmer, safer, less polluted and more walkable; foot traffic replaced car traffic and businesses became busier.
  • The same mayor has been re‑elected for 27 years (since 1999) and is described as Spain's longest‑serving mayor among large cities; the author calls Pontevedra a lively 'hidden gem' and says public space works better when prioritized for people over parked cars.

Pontevedra transformed its historic center beginning in 1999 when a new mayor removed street parking, eliminated most through‑traffic, and limited car access to residents, deliveries, taxis and emergencies. After initial fury from shop owners the city became calmer, safer, less polluted and more walkable with higher foot traffic and busier businesses. The mayor still serves 27 years later; the author calls it a lively hidden gem and argues cities improve when public space prioritizes people over car storage.

By @pitdesi
11 Twitter/X 2026-06-24 1 min read
Open

Nori's SuperNori is presented as the first AI that 'actually runs a household' —…

Why it matters

Nori's SuperNori is presented as the first AI that 'actually runs a household' — it auto-restocks groceries, books travel, plans weekly meals around family nutrition goals, and controls smart-home devices, acting proactively without being asked.

  • @heyshrutimishra claims the real competitive moat in consumer AI is 'who owns the daily routine,' not the underlying model, positioning SuperNori as a routines-first product.
  • Isaac (@IsaacDrgn) labels SuperNori the first 'Proactive Family AI Agent' built for the family caretaker; the demo and thread were posted June 24, 2026 (links to video/demo included).

SuperNori, from startup Nori, is showcased (June 24, 2026) as a proactive Family AI agent that 'runs a household' by auto-restocking groceries, booking travel, planning weekly meals to meet family nutrition goals, and controlling smart-home devices without prompting. The author and a thread by Isaac argue its moat is owning daily routines rather than the model itself.

By @heyshrutimishra
12 Twitter/X 2026-06-26 1 min read
Open

Ben Horowitz (a16z) said in a January show appearance that California's proposed…

Why it matters

Ben Horowitz (a16z) said in a January show appearance that California's proposed wealth tax is "the best strategy" he's seen to dismantle Silicon Valley's network effect.

  • Horowitz cited Norway's unrealized capital-gains tax as a parallel, claiming entrepreneurs left because private-company equity gets marked up but is illiquid — "I literally can't pay the tax...so, I have to leave the country" — and that "there are basically no tech entrepreneurs in Norway now."

Ben Horowitz of a16z warned on a January show (clip posted by @tbpn on 2026-06-26) that California's proposed wealth tax is "the best strategy" to break Silicon Valley's network effect. He compared it to Norway's unrealized capital-gains tax, claiming founders are forced to emigrate because they can't pay taxes on marked-up, illiquid private equity.

By @tbpn
13 Twitter/X 2026-06-26 1 min read
Open

@iruletheworldmo (posted 2026-06-26) states the era of public access to…

Why it matters

@iruletheworldmo (posted 2026-06-26) states the era of public access to bleeding‑edge models is over, warning that 'if these models are being taken away, we’re on the steepest part of the curve' and that access will become an 'ever receding point.'

  • @iruletheworldmo claims newer versions — mythos 5.1 and 'gpt 5.7' — are as significant a jump as mythos, argues this loss of incremental access is 'terrible for society and safety,' invokes 'Sam's' philosophy of democratizing AI, and says only getting a future 'gpt 10' would be 'an absolute societal disaster.'

Author @iruletheworldmo warns the era of public access to frontier AI is ending, alleging model availability is being pulled back and we're on the 'steepest part of the curve' (post dated 2026-06-26). They claim mythos 5.1 and gpt 5.7 match mythos's jump, argue reduced access undermines safety and Sam's democratizing philosophy, and say only seeing a future 'gpt 10' would be a societal disaster.

By @iruletheworldmo
14 Twitter/X 2026-06-26 1 min read
Open

Joe Hudson, who coaches OpenAI's research team and advises Sam Altman and leaders…

Why it matters

Joe Hudson, who coaches OpenAI's research team and advises Sam Altman and leaders at Apple and Google, is credited in Lenny Rachitsky's post (tweeted 2026-06-26) with identifying 'emotional clarity' as the key predictor of success in AI-forward teams.

  • Hudson contrasts emotional clarity with knowledge and hours—skills he says AI outperforms humans at—and defines it as the ability to stay in difficult conversations, avoid turning on oneself or others, and persist through failure.
  • He presents a four-part 'wisdom stack' to build emotional clarity: 1) Discernment, 2) 'In conflict we trust', 3) Willingness to fail, 4) Positive self-talk.

Joe Hudson, who coaches OpenAI’s research team and leaders at Apple, Google and Sam Altman, argues that “emotional clarity” — the ability to feel emotions without being driven by them — most predicts success in AI-forward work. He offers a four-part “wisdom stack” (discernment; “in conflict we trust”; willingness to fail; positive self-talk) and Lenny Rachitsky promoted the guest post on 2026-06-26.

By @lennysan
15 Twitter/X 2026-06-26 1 min read
Open

On 2026-06-26, Jeremy Allaire (@jerallaire) announced Circle's support for…

Why it matters

On 2026-06-26, Jeremy Allaire (@jerallaire) announced Circle's support for Proof's launch of x401, an open protocol to verify who authorized AI agents' actions in the so-called 'agentic economy'.

  • Proof positions x401 as answering 'who authorized the action' and pairs it with x402 (which answers 'how an agent pays'); Proof released live demos and a CLI for x401 available now.
  • Proof developed x401 with contributors across payments, identity, and AI, and Circle frames the announcement around the need for open standards for both payment and identity in the agentic economy.

Circle (Jeremy Allaire) announced on 2026-06-26 support for Proof's x401, an open protocol that verifies the authority behind AI agents' actions in the 'agentic economy.' Proof—working with payments, identity and AI contributors—positions x401 alongside x402 (payments); x401 has live demos and a CLI available from proof.com.

By @jerallaire
16 ArXiv 2026-06-25 1 min read
Open

Fast algorithms for learning a Gaussian under halfspace truncation with optimal sample complexity

Why it matters

For any ε>0 and dimension d the authors give an efficient algorithm that learns a Gaussian truncated by an unknown halfspace using n = Õ(d^2/ε^2) samples and with runtime dominated by computing the empirical covariance matrix (paper claims this is optimal in d and ε).

  • Key technical idea is a novel reinterpretation of low-degree moments via a relative truncation parameter that uniquely determines the untruncated Gaussian, enabling direct parameter recovery and avoiding the projected stochastic gradient descent used by Lee, Mehrotra & Zampetakis (FOCS'24). Paper (88 pages) accepted to COLT 2026; posted 2026-06-25.

The paper addresses learning a high-dimensional Gaussian truncated to an unknown halfspace and introduces a reinterpretation of low-degree moments in terms of a relative truncation parameter that uniquely identifies the underlying Gaussian. Using this insight the authors give an algorithm with sample complexity n = Õ(d^2/ε^2) and runtime dominated by empirical covariance computation, matching optimal bounds and removing the need for projected SGD; accepted to COLT 2026.

Authors: Haitong Liu, Deepak Narayanan Sridharan, David Steurer...
17 ArXiv 2026-06-25 1 min read
Open

Autoregressive Boltzmann Generators

Why it matters

ArBG introduces an autoregressive Boltzmann Generator framework that removes normalizing-flow topological constraints, enables sequential inference-time interventions, and scales via LLM-style architectures.

  • ArBG outperforms flow-based Boltzmann Generators across benchmarks—especially on the 10-residue peptide Chignolin—and Robin, a 132M-parameter ArBG model, reduces zero-shot energy error E-W2 on 8-residue systems by over 60%.
  • Code and models are available at https://github.com/danyalrehman/autobg; the paper appeared as an ICML 2026 (Spotlight) submission and is on arXiv: https://arxiv.org/abs/2606.27361v1.

Autoregressive Boltzmann Generators (ArBG) introduce an autoregressive generative framework for sampling molecular Boltzmann distributions that avoids normalizing-flow topology constraints and supports sequential inference-time interventions. By leveraging LLM-style scalable architectures, ArBG outperforms flow-based Boltzmann Generators on all benchmarks—most notably on the 10-residue Chignolin—and yields Robin, a 132M-parameter transferable model that cuts zero-shot E-W2 error on 8-residue systems by over 60%.

Authors: Danyal Rehman, Charlie B. Tan, Yoshua Bengio...
18 ArXiv 2026-06-25 1 min read
Open

All you need is log

Why it matters

The paper proves that any functional of W-tuples of distributions that is data-processing monotone and additive on independent products equals a positive integral of multi-way coincidence divergences C_α(π_1,...,π_W) := -log ∫ π_1^{α_1}⋯π_W^{α_W} with ∑_k α_k = 1; the α-parameter space has four necessary strata: simplex interior, mixed-sign exponent cones, a tropical boundary (max-divergences), and pairwise KL edges at simplex vertices.

  • The family is shown to be canonical via five independent derivations — structural axioms; Kolmogorov–Nagumo means with Rényi entropy axioms; classical entropy characterizations; multi-hypothesis testing error exponents; and a multi-lottery betting interpretation — and reduces to standard Rényi divergences in the two-prior case.

The paper characterizes the canonical multi-distribution generalization of Rényi divergences: any divergence on W-tuples that is monotone under data processing and additive on independent products is a positive integral over multi-way coincidence divergences Cα = -log ∫ ∏k πk^{αk} (with ∑α_k=1). The α-space decomposes into four essential strata (simplex interior, mixed-sign cones, a tropical max-divergence boundary, and pairwise KL edges). Five independent routes to the same family and a worked W=3 example with numerical checks support the claim that this is the natural multi-distribution Rényi calculus. (Akshay Balsubramani; arXiv:2606.27349v1; 2026-06-25.)

Authors: Akshay Balsubramani
19 Twitter/X 2026-06-26 1 min read
Open

@tszzl claims the popular critique that an unofficial AI licensing regime is…

Why it matters

@tszzl claims the popular critique that an unofficial AI licensing regime is 'slowing down innovation' ignores how quickly AI is moving; the Mythos incident may have accelerated oversight but such intervention was inevitable and 'earlier is better' amid exponential growth (posted 2026-06-26).

  • He views federal attention as positive: 'models being publicly delayed by a week here or there is really not the end of the world,' though he admits current procedures are 'not the right way' and expects authorities to 'figure it out.'
  • He warns it would be 'very sad' if non‑Americans are permanently left behind and urges maintaining a 'pax technologica' led by the free world (and later the unfree world) to keep the frontier accessible.

Author @tszzl argues that claims the unofficial AI licensing regime is stifling innovation ignore how rapidly AI is advancing. The Mythos episode hastened oversight but was inevitable — earlier intervention beats waiting in an exponential curve. He welcomes federal recognition of the technology’s gravity, calls short public delays tolerable, and warns against leaving non‑Americans behind.

By @tszzl
20 Twitter/X 2026-06-26 1 min read
Open

On 2026-06-26 Sarvesh Shrivastava (@bloggersarvesh) posted a playbook claiming he…

Why it matters

On 2026-06-26 Sarvesh Shrivastava (@bloggersarvesh) posted a playbook claiming he could reach $100k/month in 90 days using Claude + SEO if he woke up bankrupt.

  • He instructs users to “open Claude, paste these 20 prompts,” arguing you can do the work of an entire SEO agency while paying Claude’s $20/month instead of $10k and outrank competitors.
  • Shrivastava states he has 14 years of local SEO experience and built this playbook for himself (not for clients or as a course).

Sarvesh Shrivastava (@bloggersarvesh) claims a personal playbook that uses Claude plus 20 prompts can replace an SEO agency: pay $20/month versus $10k, execute in 90 days, and scale to $100k/month. Posted 2026-06-26, he emphasizes this is his contingency plan, built from 14 years of local SEO experience and intended for his own use.

By @bloggersarvesh
21 ArXiv 2026-06-25 1 min read
Open

BOWConnect (Raxit et al., accepted to IROS 2026, published 2026-06-25) integrates…

Why it matters

BOWConnect (Raxit et al., accepted to IROS 2026, published 2026-06-25) integrates Bayesian Optimization over Windows (BOW) as a learned steering function inside a bidirectional parallel kinodynamic planner to learn local cost maps and guide constraint-aware control sampling.

  • Evaluated on ten benchmark environments, BOWConnect achieved a 100% success rate and delivered the fastest or near-fastest planning times in narrow-passage and non-convex scenarios; real-world tests on a ground vehicle and a quadrotor ran in real time with no collisions.

BOWConnect is a bidirectional, parallel kinodynamic motion planner that uses Bayesian Optimization over Windows (BOW) as a learned steering function to build local cost maps guiding control sampling. It addresses sample inefficiency, poor dynamic heuristics, and narrow passages via parallel trees, spatial hashing for fast connections, and a boundary-value solver. On ten benchmarks it achieved 100% success and real-time, collision-free deployment on ground and aerial robots.

Authors: Sourav Raxit, Abdullah Al Redwan Newaz, Jose Fuentes...
22 ArXiv 2026-06-25 1 min read
Open

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

Why it matters

LA4VLA constructs LA4-33K, a dataset of 33,000 Language-Action (LA) episodes by decomposing expert demonstration trajectories into atomic action segments paired with low-level action descriptions, created without additional robot data collection.

  • The authors train a lightweight LA4VLA-1B (1 billion parameters) VLA model and compare three pretraining paradigms (LA-only, sequential LA→VLA, mixed LA+VLA); mixed LA-VLA pretraining raises average success rates versus no-pretraining by up to 17.8 percentage points in simulation and 45.0 percentage points in real-world tasks.

LA4VLA proposes language-action pretraining to teach language-conditioned action priors without visual inputs, addressing VLA policies' tendency to rely on visual shortcuts when visual-action signals dominate sparse language-action supervision. The method decomposes demonstrations into atomic action segments with low-level action descriptions (LA4-33K) and trains a 1B-parameter model (LA4VLA-1B). Across simulation and real robot tests, LA-pretraining outperforms matched VLA pretraining and mixed LA+VLA yields the largest gains (up to +17.8 pp sim, +45.0 pp real). Summary based on the paper abstract; full text was not reviewed.

Authors: Tao Lin, Yuxin Du, Yiran Mao...
23 ArXiv 2026-06-25 1 min read
Open

RouterVLA: Turning Smoke Tests into Supervision for Heterogeneous VLA Selection

Why it matters

RouterVLA uses outcome-disjoint cross-fitting on 34,752 LIBERO-Plus rollout records to build probe profiles for frozen VLA experts and raises held-out success from 0.4686 to 0.6149 (a +14.64 percentage-point gain) using a transparent probe-success rule.

  • Under the scalar-only profiles studied, learned scorers are statistically indistinguishable from the simple rule; reusing the scored trial inflates measured gain by 1.87×, indicating routing/commissioning drives system gains beyond mere model scaling.

RouterVLA evaluates whether pre-deployment smoke-test rollouts can supervise selection among heterogeneous vision–language–action (VLA) policies by using outcome-disjoint cross-fitting: one set of probes builds frozen-expert profiles while separate scored trials estimate held-out performance. On 34,752 LIBERO-Plus rollouts the approach raises held-out success from 0.4686 to 0.6149 (+14.64 pp). Learned scalar scorers matched a transparent probe-success rule, and reusing the scored trial overstates benefits by 1.87×. Full text was not available; results suggest commissioning-aware routing adds system-level value beyond per-model scaling.

Authors: Xingyu Ren, Chugang Yi, Ge Ma...
24 Twitter/X 2026-06-26 1 min read
Open

On 2026-06-26 @cryptopunk7213 claims OpenAI's new GPT-5.6 'Sol' has officially…

Why it matters

On 2026-06-26 @cryptopunk7213 claims OpenAI's new GPT-5.6 'Sol' has officially taken the #1 spot and "beats Mythos" on coding tasks.

  • The author reports GPT-5.6 "crushed" the Terminal 2.1 benchmark while using one‑third the tokens; OpenAI previewed GPT-5.6 Sol (limited preview), plus Terra and Luna, and says Sol includes an "ultra mode" that spins up multiple sub‑agents and will have a limited public release in a few weeks.

Cryptopunk7213 celebrates OpenAI's June 26, 2026 preview of GPT-5.6, claiming Sol has taken the #1 spot over Mythos and "crushed" the Terminal 2.1 benchmark while using one‑third the tokens. OpenAI's announcement previews Sol (limited), Terra, and Luna; Sol includes an "ultra mode" that spawns sub‑agents and will see a limited public release in a few weeks.

By @cryptopunk7213
25 ArXiv 2026-06-25 1 min read
Open

Data-Driven Duration Management -- Term Structure Forecasting Using Machine Learning

Why it matters

Neural networks consistently beat classical term‑structure methods (Dynamic Nelson‑Siegel, PCA) on both U.S. Treasury and ECB zero‑coupon bond forecasts and in downstream portfolio performance; evaluation used RMSE, MAE, directional accuracy plus an economic bond‑trading metric (Lausser et al., 2026‑06‑25).

  • Best models differ by market: for the U.S. a direct‑forecasting NN that uses DNS factors for zero‑rate dimensionality reduction and an Autoencoder to extract macroeconomic features performed best; for Europe a factor‑based NN using PCA‑derived zero‑rate factors without macro integration was optimal.

The paper evaluates forecasting of U.S. and European zero‑coupon yield curves by comparing classical approaches (Dynamic Nelson‑Siegel, PCA) with multiple neural‑network architectures, incorporating macroeconomic inputs and Autoencoders. Using statistical metrics (RMSE, MAE, directional accuracy) plus a bond‑trading performance test, the authors find NNs improve both forecast accuracy and portfolio returns; optimal architectures differ across the two markets.

Authors: Tobias Lausser, Joao Eduardo Vuolo, Rudi Zagst
26 ArXiv 2026-06-25 1 min read
Open

Asymptotically Optimal Learning for Parametric Prophet Inequalities

Why it matters

Characterized the optimal full-information asymptotic competitive ratio for i.i.d. rewards from an exponential-type parametric family (unknown θ): for unbounded-support distributions the limit equals ((θ/(θ - c_+))^{c_+/θ}) / Γ(1 - c_+/θ), while for bounded-support power-family the limit is 1.

  • Propose a confidence-based dynamic-programming online learning policy that, using only online observations and no external offline samples, asymptotically attains the same optimal competitive ratio as the full-information benchmark.
  • Derive distribution-specific convergence rates for canonical examples (including exponential and Pareto) and validate the algorithm with synthetic numerical experiments; authors Jung-hun Kim, Anna Grebennikova, Vianney Perchet, arXiv:2606.26893v1 (2026-06-25).

The paper studies online learning for prophet inequalities when rewards are i.i.d. from an exponential-type parametric family with unknown θ (includes exponential, Pareto, bounded-support power-family). It gives a closed-form optimal full-information asymptotic competitive ratio (unbounded case: ((θ/(θ−c_+))^{c_+/θ})/Γ(1−c_+/θ); bounded case: 1) and designs a confidence-based dynamic-programming policy that, from only online samples, achieves these limits with distribution-specific convergence rates and synthetic validation.

Authors: Jung-hun Kim, Anna Grebennikova, Vianney Perchet
27 ArXiv 2026-06-25 1 min read
Open

XMSE-Aware Adaptive Empirical Bayes Estimation

Why it matters

Proposes an "XMSE-aware mixed estimator" (2026-06-25, Chen & Zheng) that linearly interpolates between maximum likelihood (ML) and a kernel-based empirical Bayes (EB) estimator; the fixed-weight excess mean squared error (XMSE) is a scalar quadratic, yielding a closed-form oracle mixing weight that is provably no worse than both ML and the base EB at the XMSE scale.

  • Provides a plug-in implementation using finite-sample XMSE approximations that is consistent and attains a second-order oracle regret rate when the oracle weight is interior; theoretical extensions include transferring the regret bound to the fixed-weight risk curve, a thresholded boundary rule, compact kernel families, and finite/growing kernel dictionaries with high-probability oracle bounds.
  • Empirical validation on finite-impulse-response simulations and public benchmarks (Silverbox, Cascaded Tanks) against SURE-tuned, hard-selection, and trace-corrected baselines shows the estimator preserves regularization benefits when kernels are well-aligned and retreats toward ML under kernel misspecification.

XMSE-Aware Adaptive Empirical Bayes introduces a mixed estimator that interpolates between ML and kernel EB shrinkage to control excess mean squared error (XMSE). Using a fixed-weight XMSE that is a scalar quadratic, the authors derive a closed-form oracle mixing weight and a consistent plug-in implementation with a second-order oracle regret rate. Theory covers kernel-family extensions; experiments on FIR simulations and Silverbox/Cascaded Tanks benchmarks demonstrate robustness to kernel misspecification.

Authors: Minghao Chen, Jiale Zheng
28 ArXiv 2026-06-25 1 min read
Open

Beyond Global Divergences: A Local-Mass Perspective on Bayesian Inference

Why it matters

Introduces two tools: Mass Index (records polynomial and logarithmic decay scales of local mass) and regularised extended KL (RE-KL), a set-localised divergence that admits singular components.

  • Mass Index shows how Bayesian updating alters local mass: power-log likelihood factors shift local-mass scales explicitly, while parameter-dependent supports or their smooth softenings change the local decay scale by varying the mass remaining near a parameter.
  • Using local RE-KL the authors prove absolute, relative, and directional inequalities for comparing local small-ball masses under the two KL directions; paper is 28 pages (3 figures, 2 tables), posted to arXiv:2606.27090v1 on 2026-06-25, code at https://github.com/Forsythia0604/Local-Mass-Framework.

Beyond Global Divergences develops a local-mass framework for Bayesian inference, introducing the Mass Index and regularised extended KL (RE-KL) to quantify polynomial/logarithmic decay of local mass and set-localised divergences (handling singular supports). The authors prove absolute, relative and directional inequalities comparing small-ball masses under forward/reverse KL, present controlled experiments, and release code.

Authors: Hanli Xu, Fengxiang He, Sarat Moka
29 ArXiv 2026-06-25 1 min read
Open

SAM2Matting: Generalized Image and Video Matting

Why it matters

SAM2Matting (Ruiqi Shen, Guangquan Jie, Chang Liu, Henghui Ding; arXiv 2026-06-25; ECCV 2026 extended) is a tracker-to-matting framework that augments foundational VOS trackers (e.g., SAM2, SAM3) with a region-proposal bridge and dedicated matting heads, decoupling temporal tracking from fine-grained matting.

  • Despite being trained only on images, SAM2Matting claims new state-of-the-art video-matting performance, supports diverse prompt types, maintains strong temporal consistency, and generalizes across human-centric and in-the-wild scenarios.

SAM2Matting reframes video matting as a tracker-to-matting pipeline: a high-fidelity matting module plus region-proposal bridge built on SAM-family trackers preserves temporal robustness while matting heads recover fine detail. Despite image-only training, authors report SOTA video-matting performance, diverse prompt support, and strong cross-domain generalization. Only the paper's abstract was available; full-text evaluation details are not provided.

Authors: Ruiqi Shen, Guangquan Jie, Chang Liu...
30 ArXiv 2026-06-25 1 min read
Open

RayPE: Ray-Space Positional Encoding for 3D-Aware Video Generation

Why it matters

RayPE injects per-token 6D Plucker coordinates additively into queries and keys of self-attention (with a query/key flip) so the symmetric identity matches the Plucker reciprocal product; the resulting attention score cleanly decomposes into a content term, a geometry term, and two cross-terms, each found necessary by experiments.

  • To stabilize across heterogeneous camera-translation scales, RayPE decouples ray direction from moment magnitude, gates the encoding by a learned function of the log-magnitude, and applies RMSNorm to align with QKNorm-normalized content; the module is zero-initialized, adds <0.1% parameters to a pretrained video DiT, and improves camera controllability, cross-frame 3D consistency, and overall video quality on a four-dataset training mixture.

RayPE augments video diffusion transformers with 3D-aware positional encoding by injecting per-token 6D Plucker ray coordinates into queries and keys (with a query/key flip) so attention bilinearly captures the Plucker reciprocal product. The additive design yields separable content, geometry, and cross-terms; stability is achieved via direction/magnitude decoupling, log-magnitude gating, and RMSNorm. The zero‑init module adds <0.1% params to a pretrained video DiT and improves camera control and 3D consistency on a four-dataset mixture.

Authors: Minghao Yin, Jiahao Lu, Wenbo Hu...
31 ArXiv 2026-06-25 1 min read
Open

PhysiFormer: Learning to Simulate Mechanics in World Space

Why it matters

PhysiFormer (Yiming Chen, Yushi Lan, Andrea Vedaldi; ArXiv 2026-06-25) is a diffusion transformer that samples future 3D mesh vertex trajectories in world coordinates from initial vertex positions, velocities, and material type (rigid or elastic) using a denoising diffusion process directly in coordinate space.

  • The model was trained on over 100k simulated trajectories and uses attention factorised over time, space, and objects for efficiency, enabling permutation-invariant multi-object reasoning and generalisation to mixed-material settings, unseen real-world geometries, and larger object counts.
  • PhysiFormer captures uncertainty to produce diverse plausible futures and substantially outperforms autoregressive baselines on trajectory accuracy, rigidity preservation, and momentum-based physical consistency.

PhysiFormer predicts physically-plausible 3D object motion by running a denoising diffusion process on mesh vertex coordinates in world space, conditioned on initial vertex positions, velocities, and material type (rigid/elastic). Trained on >100k simulated trajectories and using factorised attention for time/space/objects, it yields permutation-invariant multi-object reasoning, diverse stochastic futures, and better accuracy, rigidity, and momentum consistency than autoregressive baselines.

Authors: Yiming Chen, Yushi Lan, Andrea Vedaldi
32 ArXiv 2026-06-25 1 min read
Open

DnA: Denoising Attention for Visual Tasks

Why it matters

DnA (Denoising Attention) uses a positive query to select class-relevant image features and a negative query to select closely associated but irrelevant features, then projects their interactions into two distinct subspaces with larger principal angles to promote subspace separation and improved discriminability.

  • With a ViT-B backbone DnA yields an absolute +0.8% top-1 on ImageNet‑1K; it also improves video understanding by +1.8% for video transformers and +0.5% for video LLMs.
  • The authors report extensive empirical analyses that justify the two interacting-subspace design and the claimed denoising effect versus standard softmax attention, which they identify as producing noisy attention patterns.

DnA (Denoising Attention) modifies multihead attention by adding positive and negative queries and projecting their interactions into two subspaces with larger principal angles to separate relevant from correlated-but-irrelevant features. Evaluated with a ViT‑B backbone, DnA improves ImageNet‑1K by 0.8% and yields gains on video tasks (+1.8% video transformers, +0.5% video LLMs); extensive experiments are reported to support design choices.

Authors: Ron Campos, Subhajit Maity, Xin Li...
33 ArXiv 2026-06-25 1 min read
Open

World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays

Why it matters

Recurrent Generative Replay (REGEN) leverages World Action Models (WAMs) to synthesize pseudo-replay trajectories by recursively querying a generative world+action model conditioned on prior task instructions and current-task observations; evaluated in both simulation and real-world robot manipulation (Govind et al., arXiv 2026-06-25).

  • REGEN reduces catastrophic forgetting by up to 50% relative to sequential fine-tuning and approaches the performance of privileged experience-replay methods that require access to stored real demonstration data.
  • The authors identify long-horizon visual degradation and action–observation inconsistency as the primary bottlenecks limiting the fidelity and effectiveness of generated replay trajectories.

World Action Models (WAMs) are used to generate future visual observations and underpin Recurrent Generative Replay (REGEN), a continual imitation-learning method that synthesizes pseudo-replay trajectories conditioned on prior task instructions and current observations. Evaluated in simulation and real-world manipulation, REGEN cuts catastrophic forgetting by up to 50% versus sequential fine-tuning and nears privileged experience-replay performance; primary failure modes are long-horizon visual degradation and action–observation inconsistency.

Authors: Manish Kumar Govind, Dominick Reilly, Smit Patel...
34 ArXiv 2026-06-25 1 min read
Open

Bridging Talk and Thought: Understanding Dialogue Dynamics Across Collaborative Problem-Solving Contexts

Why it matters

Zhengyuan Liu, Stella Xin Yin, Min-Yen Kan, and Nancy F. Chen (published 2026-06-25) introduce a hierarchical two-layer coding scheme that integrates cognitive and non-cognitive problem solving with explicit metacognitive regulatory mechanisms.

  • They validate the framework across nine datasets spanning multiple domains and report that metacognitive regulation reliably discriminates deeper forms of human–AI and multi-agent collaboration.

The authors present a conceptual framework and hierarchical two-layer coding scheme to analyze dialogue during collaborative problem solving, integrating cognitive, non-cognitive, and metacognitive regulatory processes. Applied to nine cross-domain datasets, the approach uncovers how humans and autonomous agents coordinate knowledge and effort, showing metacognitive regulation as a key marker of deeper collaboration. Full text was not available for review; summary is based on the abstract.

Authors: Zhengyuan Liu, Stella Xin Yin, Min-Yen Kan...
35 ArXiv 2026-06-25 1 min read
Open

LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank

Why it matters

LLM-based pipeline achieved up to 91% document-level precision for eligibility decisions, operating conservatively to minimize false acceptance (paper reports high precision at the document level).

  • This is the first case study applying generative LLM information-extraction to the German Central Bank’s collateral-eligibility checks, using a three-stage pipeline (extraction, normalization, interpretation) to handle noisy, bilingual (German–English) prospectuses.
  • Authors introduce a value-based evaluation with an LLM-as-a-judge for semantic assessment, contrasting prior span-based NER approaches that struggle with OCR noise, linguistic variance, and expensive manual annotation.

The paper addresses automated verification of securities collateral eligibility at the German Central Bank by replacing rigid span-based NER with a generative LLM information-extraction pipeline (extraction, normalization, interpretation) tailored for noisy, semi-structured, bilingual prospectuses. Presented as the first case study in this domain, the approach attains up to 91% document-level precision and introduces a value-based LLM-as-judge evaluation that better captures semantic correctness than location-based metrics.

Authors: Serhii Hamotskyi, Akash Kumar Gautam, Christian Hänig
36 ArXiv 2026-06-25 1 min read
Open

Mapping Political-Elite Networks in Europe with a Multilingual Joint Entity-Relation Extraction Pipeline

Why it matters

Modular open-weight multilingual pipeline that builds signed, temporal knowledge graphs from news using span-based NER, a three-stage linking cascade mapping mentions to language-independent Wikidata IDs, and an ontology-constrained mixture-of-experts model with guided decoding.

  • Evaluation: full-coverage spot-check on a 3,491-relation gold standard reports textual correctness 68.2% strict and 93.7% lenient.
  • Case studies: Austria — reconstructs a political party’s lifecycle, dating internal fractures and tracking personnel into successor factions and court convictions; Poland — uncovers overlapping state-enterprise patronage networks and a structurally balanced signed conflict network between PO (Platforma Obywatelska) and PiS (Prawo i Sprawiedliwość).

Solovev and Lasser (2026) introduce an open-weight, multilingual pipeline that jointly extracts entities and signed, directed relations from news to build temporal knowledge graphs. It combines span-based NER, a three-stage Wikidata linking cascade, and an ontology-constrained mixture-of-experts with guided decoding. A 3,491-relation spot-check yields 68.2% strict (93.7% lenient) correctness; Austrian and Polish case studies validate political-network recovery.

Authors: Kirill Solovev, Jana Lasser
37 ArXiv 2026-06-25 1 min read
Open

Blackwell Approachability and Gradient Equilibrium are Equivalent

Why it matters

Proves algorithmic equivalence between Gradient Equilibrium (GEQ) and Blackwell approachability: each can be solved using a black-box oracle for the other with no asymptotic loss in the oracle's error rate, and the reductions are efficient.

  • By combining this with known approachability–regret–calibration equivalences, GEQ is therefore algorithmically equivalent to regret minimization and calibration; the reductions preserve refined guarantees such as optimism and strong adaptivity.
  • Gives necessary and sufficient conditions for GEQ and reductions between unconstrained and constrained decision sets; paper (Brian W. Lee, Nika Haghtalab, Michael I. Jordan, Ryan J. Tibshirani) is 30 pages, on arXiv (2606.27315v1) and accepted to COLT 2026 (posted 2026-06-25).

Gradient equilibrium (GEQ) is shown equivalent to Blackwell approachability: the authors provide mutual, efficient reductions that use a GEQ black-box to solve approachability problems (and vice versa) with no asymptotic loss in error. Combined with known approachability–regret–calibration equivalences, GEQ is thus algorithmically equivalent to regret minimization and calibration; they also state necessary/sufficient conditions and reductions for constrained vs. unconstrained GEQ.

Authors: Brian W. Lee, Nika Haghtalab, Michael I. Jordan...
38 ArXiv 2026-06-25 1 min read
Open

DanceOPD: On-Policy Generative Field Distillation

Why it matters

DanceOPD (2026-06-25) introduces on-policy generative field distillation for flow-matching models: each sample is routed to one capability-specific velocity field, the student queries a low-noise student-induced state, and training uses a simple velocity MSE objective to learn composition of capabilities.

  • Capabilities (text-to-image, local editing, global editing) are represented as velocity fields over a shared flow state; the student learns from fields queried on its own rollout states and can absorb operator-defined fields such as classifier-free guidance (CFG), improving multi-capability composition while preserving anchor generation quality.
  • Technical report (39 pages, 13 figures, 9 tables) with experiments on T2I, editing, realism-field absorption, and CFG absorption that demonstrate strengthened target capabilities and practical applicability; project page: https://danceopd.github.io/.

DanceOPD tackles the conflict between text-to-image, local editing, and global editing by framing each capability as a velocity field in a shared flow-matching state space. The method routes samples to a single capability field, queries a low-noise student-induced state, and trains the student on its own rollouts with a velocity MSE loss. Experiments report improved multi-capability composition and preservation of anchor generation quality, and the method can absorb operator-defined fields such as classifier-free guidance.

Authors: Wei Zhou, Xiongwei Zhu, Zelin Xu...
39 ArXiv 2026-06-25 1 min read
Open

Simulation-based inference for rapid Bayesian parameter estimation in epidemiological models: a comparison with MCMC

Why it matters

SBI (neural posterior estimation) calibrated a mechanistic SECIR COVID‑19 model on Germany 2020 ICU-occupancy data: for 31-day inference windows SBI matched MCMC posteriors while running ~60–70 seconds on a single GPU versus ~1,000 seconds for MCMC (CPU); for a 201-day reconstruction SBI averaged ~157 seconds vs >19,000 seconds for MCMC.

  • Posterior agreement was evaluated with Wasserstein distances, Kullback–Leibler divergences, and posterior predictive checks; SBI accurately reproduced observed ICU trajectories across 31-day windows and preserved the dominant posterior structure in the more uncertain 201-day problem.
  • SBI leverages neural posterior estimation and combined CPU+GPU resources to provide a rapid, scalable Bayesian calibration alternative to standard MCMC for high-dimensional, nonlinear epidemiological models.

Simulation-based inference (SBI) using neural posterior estimation is evaluated as a fast alternative to MCMC for Bayesian calibration of a mechanistic SECIR model to COVID‑19 ICU occupancy data from Germany (2020). Across 31-day windows SBI recovered MCMC-like posteriors and trajectories; on a challenging 201-day reconstruction SBI preserved key posterior structure while reducing compute from >19,000s to ~157s.

Authors: Alina Bazarova, Johann Fredrik Jadebeck, Henrik Zunker...