Briefing · 2026-07-17

Your briefing

90 ranked ·

Today's dispatch

Filed · 90 ranked

  1. 93 score ArXiv · Must read · 1 min T^2MLR: Transformer with Temporal Middle-Layer Recurrence Paper (Ziyang Cai et al., published on arXiv 2026-07-16) introduces T^2MLR (Transformer with Temporal Middle-Layer Recurrence) that fuses a cached middle-layer representation from the previous token into an earlier layer of the current token to preserve intermediate reasoning states.
  2. 84 score ArXiv · Worth reading · 1 min Pretraining Data Can Be Poisoned through Computational Propaganda Graf et al. (Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith, David Kohlbrenner, Kyle Lo; arXiv 2026-07-16) show that pretraining-data poisoning is feasible via public web discussion interfaces, extending beyond prior Wikipedia‑focused attacks and identifying third‑party webpage content as a viable attack vector.
  3. 78 score ArXiv · Worth reading · 1 min BadWAM: When World-Action Models Dream Right but Act Wrong BadWAM defines World-Action Drift Attacks against World-Action Models (WAMs), characterizing the attack surface along two criteria—attack strength and stealthiness—and instantiates two attacks: an action-only attack and an imagination-preserving attack.
  4. 78 score ArXiv · Worth reading · 1 min Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search Using a ReAct agent over HotpotQA, the authors replayed 1,000 development questions and computed a Counterfactual Trajectory Utility (CTU) by deleting each read document and re-running the trajectory; across 23,322 document observations CTU and Static RAG Utility (SRU) are nearly independent (Spearman rho = -0.026).
  5. 78 score ArXiv · Worth reading · 1 min Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models LLMs often violate the law of total probability under test: using binary-tree partitions (prompting models with verbalized subpopulation descriptions and aggregating estimates) yields inconsistent population-level marginals across partitions for state-of-the-art 'frontier' models and multiple problem domains.
  6. 78 score ArXiv · Worth reading · 1 min Sharp Stability Threshold and Certification for Designing Stable Residual Architectures Proposes the sublinear-growth principle and proves a sharp stability threshold q = 1 for residual-block velocity fields obeying ‖v(x,t)‖ ≤ c‖x‖^q + b (q ∈ [0,1]); classical ODE theory gives global forward flow for q ≤ 1 and divergent velocity fields for any q > 1.
  7. 78 score Twitter/X · Worth reading · 1 min On 2026-06-04 Andrew Ng announced a short deeplearning.ai course (built with Red… On 2026-06-04 Andrew Ng announced a short deeplearning.ai course (built with Red Hat) taught by Cedric Clyburn on serving LLMs to many concurrent users with low latency and reasonable cost.
  8. 77 score ArXiv · Worth reading · 1 min In-Place Tokenizer Expansion for Pre-trained LLMs Introduces in-place tokenizer expansion: continue the source tokenizer's BPE merges on a multilingual corpus, copy carried-over embedding rows unchanged, initialize new token embeddings as the mean of their source sub-token embeddings, and use a two-stage adaptation (embedding-only training, then full-model continued pre-training) to recover source-checkpoint quality.
  9. 76 score Twitter/X · Worth reading · 8 min On 2026-05-31 Richard S. On 2026-05-31 Richard S. Sutton asserts that generative AI trained by supervised learning (LLMs, image/video models, world models) cannot make novel discoveries because it lacks runtime Evaluation and thus cannot perform selective retention — stochastic generation yields novelty but not evaluated, retained discoveries.
  10. 76 score Twitter Article · Worth reading · 8 min T0nyav’s 11 April 2026 post frames Anthropic’s Mythos and Project Glasswing as… T0nyav’s 11 April 2026 post frames Anthropic’s Mythos and Project Glasswing as evidence the internet’s permissionless frontier is closing: Mythos won’t be generally released and access appears concentrated among enterprise partners (AWS, Apple, Cisco, CrowdStrike, Google, Microsoft, NVIDIA, etc.). Citing Frederick Jackson Turner (1893), Rudolf…
  11. 75 score Twitter Article · Worth reading · 1 min Mintlify's AI assistant now uses ChromaFS, a virtual filesystem that turns UNIX… Mintlify's AI assistant now uses ChromaFS, a virtual filesystem that turns UNIX commands into database queries to simulate a sandbox. Announced 2026-04-02 by @cdxker, this replaced chunk-based RAG (which missed cross-page context) and avoided slow sandboxes (~46s startup, $70k/yr), cutting latency to ~10ms and serving 30,000+ conversations daily.
  12. 73 score ArXiv · Worth reading · 1 min Data Driven Block Replacement Scheduling Presents Hoeffding- and Bernstein-based lower-confidence-bound bandit algorithms for choosing replacement interval k ∈ {1,…,K}, achieving O(K log T) regret and matching the Lai–Robbins lower bound; correlated variants attain O((K−k*) log T) regret and require only O(1) direct pulls of suboptimal arms k < k*.
  13. 72 score ArXiv · Worth reading · 1 min What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity The paper proves the 2025 conjecture of Patel et al. that a bounded second-order heterogeneity assumption yields improved convergence guarantees for Local SGD (Federated Averaging) on general convex objectives; the authors (K. K. Patel, R. Islamov, S. U. Stich, A. Lucchi, E. Gorbunov, L. Wang) give improved upper bounds and nearly-tight lower bounds (arXiv:2607.14731v1, published 2026-07-16).
  14. 72 score ArXiv · Worth reading · 1 min Subjective Risk Decomposition: A New View for Uncertainty Quantification Subjective-risk decomposition: Alamri, Caprio, and Brown (arXiv:2607.15196v1, published 2026-07-16) show epistemic and aleatoric uncertainty can be derived as consequences of modelling choices by decomposing a subjective risk specified by a strictly proper loss; reverse cross-entropy recovers classic information-theoretic uncertainty terms.
  15. 72 score ArXiv · Worth reading · 1 min On-Policy Delta Distillation Introduces the delta signal as a distillation reward: delta = (teacher model) − (base model before instruction tuning), designed to capture changes induced by reasoning tuning; the method is named On-Policy Delta Distillation (OPD²).
  16. 72 score ArXiv · Worth reading · 1 min SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration SearchOS reformulates open-domain information seeking as relational schema completion with grounded citations and externalizes progress into Search-Oriented Context Management (SOCM) components: Frontier Task, Evidence Graph, Coverage Map, and Failure Memory.
  17. 72 score ArXiv · Worth reading · 1 min Mask-Aware Policy Gradients for Diffusion Language Models Proposes mask-aware policy gradients for Masked Diffusion Language Models (MDLMs), formalizing generation as a two-stage action MDP (token placement then which positions to remask) and showing the policy gradient decomposes into a token term and a masking term.
  18. 72 score ArXiv · Worth reading · 1 min The Industrialization of Research ; On AI-Driven Science and Its Consequences Frames 'industrialization of research' as a shift from a craft model to an automated, supervised pipeline; cites the US Department of Energy's Genesis Mission as the most ambitious current instantiation (ArXiv:2607.15164v1, published 2026-07-16).
  19. 72 score ArXiv · Worth reading · 1 min Delocalization of bias in unadjusted Hamiltonian Monte Carlo and underdamped Langevin For unadjusted Hamiltonian Monte Carlo and underdamped Langevin, controlling the W2 bias of any K-dimensional marginal of a d-dimensional target requires O(√K) integration steps (up to log d factors) under assumptions of weak or sparse interactions.
  20. 70 score ArXiv · Worth reading · 1 min RTS Smoother-Guided Learning of Physics-Based Neural Differential Models Introduces an RTS-smoother-guided hybrid neural–physics ODE learning scheme that alternates between (1) latent-state inference with a Rauch–Tung–Striebel (RTS) smoother treating model parameters as fixed and (2) neural-network parameter updates via backpropagation on the smoothed trajectories; iterations continue until a stopping criterion.
  21. 68 score Twitter/X · Worth reading · 1 min Author @adxtyahq (published 2026-07-16) frames the prompt “design Claude Code… Author @adxtyahq (published 2026-07-16) frames the prompt “design Claude Code from scratch” as an Anthropic interview question and presents an eight-step blueprint, while noting “Probably not how Claude Code is actually built.”
  22. 68 score ArXiv · Worth reading · 1 min Online Neural Space Time Memory for Dynamic Novel View Synthesis Proposes Online Neural Space Time Memory (NST-Mem) that decouples memory update and application: periodic (infrequent) gradient-based memory updates plus per-frame memory application via cross-view attention to handle deformations between stored memory and current frames (Elmieh et al., arXiv 2026-07-16).
  23. 68 score Twitter/X · Worth reading · 2 min Saronic announced Port Alpha Saronic announced Port Alpha: a $3.2 billion greenfield shipyard in Brownsville, Texas on 835 acres (expandable to 4,400 acres), built for software-defined, robotics-enabled production and autonomy; opening in 2028 to produce 850-foot vessels initially and 1,200-foot ships later, with projected impacts of >$160 billion for Cameron County, $264.5 billion for Texas, and up to 10,000 direct jobs.
  24. 68 score ArXiv · Worth reading · 1 min Linear representations of grammaticality in neural language models Li & Kim (2026-07-16; arXiv:2607.15175v1) apply mass-mean probing and find grammatical vs. ungrammatical sentences are linearly separable in sentence representations of a wide range of pretrained neural language models; this representational separation is not fully explained by correlated sentence-level factors (e.g., lexical frequency, plausibility, world knowledge).
  25. 68 score ArXiv · Worth reading · 1 min RoboTTT: Context Scaling for Robot Policies RoboTTT scales visuomotor context to 8K timesteps (three orders of magnitude beyond prior policies) without increasing inference latency, enabling one-shot in-context imitation, on-the-fly policy improvement, robustness to perturbations, and better long-horizon multi-stage performance.
  26. 68 score ArXiv · Worth reading · 1 min MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators MeanFlowNFT adapts DiffusionNFT's forward-process RL to MeanFlow average-velocity generators by deriving an induced instantaneous-velocity predictor via a MeanFlow identity, enabling reward optimization while preserving MeanFlow's fast few-step (average-velocity) sampling.
  27. 68 score ArXiv · Worth reading · 1 min AutoSynthesis: An agentic system for automated meta-analysis AutoSynthesis is an end-to-end multi-agent system that, given a natural-language research question, automates search-strategy formulation, literature retrieval, screening, full-text eligibility assessment, quantitative-statistic extraction, standardized effect-size computation, and random-effects meta-analysis.
  28. 68 score ArXiv · Worth reading · 1 min Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents The paper builds a multi-agent framework combining Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and per-party Retrieval-Augmented Generation (RAG) to create manifesto-bound partisan LLM agents; DPO gives aggressive party-specific personas while RAG enforces manifesto grounding and MILT (Multi-Layered Information Lineage Topology) traces every clause into five provenance states.
  29. 68 score Twitter/X · Worth reading · 1 min @adxtyahq (posted 2026-07-16 19 @adxtyahq (posted 2026-07-16 19:21:53+00:00) claims Kimi K3 "feels like China just open-sourced a Fable-5 class model" and reports it generated exceptionally polished UI from a single prompt.
  30. 68 score Twitter/X · Worth reading · 5 min Kimi K3 is a 2.8 trillion-parameter native multimodal open model with 1M-token… Kimi K3 is a 2.8 trillion-parameter native multimodal open model with 1M-token context, a latent MoE that activates 16 of 896 experts, KDA attention and attention residuals, quantization-aware training (mxfp4 weights, mxfp8 activations), and claimed ~2.5x scaling efficiency over K2; pricing: $0.30/mtok cache-hit, $3.00/mtok cache-miss, $15.00/mtok output, and >90% cache-hit rate on coding workloads.
  31. 68 score ArXiv · Worth reading · 1 min Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy Benchmarked six MLLMs (three closed-source, three open-source) on the Scientific Visualization Literacy Assessment: 49 items, 18 scientific visualizations, 8 techniques, 11 task types, using a closed-world protocol and a human baseline from 485 participants (arXiv preprint published 2026-07-16).
  32. 68 score Twitter/X · Worth reading · 1 min Alex Cheema says Kimi K3 will require 4 × 512GB M3 Ultra Mac Studios (noting 2TB… Alex Cheema says Kimi K3 will require 4 × 512GB M3 Ultra Mac Studios (noting 2TB @ 3.2TB/s) and he expects it can run at 30+ tokens/sec using MTP + tensor parallelism over RDMA via Thunderbolt 5.
  33. 68 score ArXiv · Worth reading · 1 min AHEAD: Anticipatory Hand-Driven Teleoperation via Human Intent Prediction AHEAD (real-time VR teleoperation) uses a short window of 3D hand and head signals plus scene context with an attention-based classifier and a state machine in a digital twin; Top-1 intent accuracy is 76% for grasp-object prediction and 76% for target-slot prediction.
  34. 65 score Twitter/X · Worth reading · 2 min Karpathy identifies two distinct audiences Karpathy identifies two distinct audiences: (1) people who tried the free ChatGPT tier last year and judge AI by its visible quirks/hallucinations, and (2) paying technical professionals who use state-of-the-art agentic models (he names OpenAI Codex and Anthropic/Claude Code) professionally in programming, math, and research.
  35. 62 score ArXiv · Worth reading · 1 min Tamed Stochastic Gradient Hamiltonian Monte Carlo Wang and Zhang (arXiv:2607.14862v1, published 2026-07-16) propose tamed stochastic gradient Hamiltonian Monte Carlo (tSGHMC) for sampling and stochastic optimization when stochastic gradients grow superlinearly.
  36. 62 score ArXiv · Worth reading · 1 min Hierarchical Denoising For Multi-Step Visual Reasoning HDR (Hierarchical Denoising for Visual Reasoning) uses tree-structured video latents and a sparse hierarchical attention pattern (SHAP) to enable coarse-to-fine, multi-step visual reasoning; evaluated on a level-stratified benchmark covering six tasks: maze navigation, Tower of Hanoi, one-line drawing, sliding puzzle, Sokoban, and water pouring.
  37. 62 score Twitter/X · Worth reading · 2 min Andrew Ng (post published 2026-06-01) identifies the AI Forward Deployed Engineer… Andrew Ng (post published 2026-06-01) identifies the AI Forward Deployed Engineer (FDE) as a rising Silicon Valley role: engineers embedded in client organizations to build and tune agentic workflows and customize off‑the‑shelf LLMs.
  38. 62 score ArXiv · Worth reading · 1 min Stigmergic Graph Memory: An Environment-Aware Approach for Many-to-Many Multi-Agent Pickup and Delivery Stigmergic Graph Memory (SGM) is a bounded, decaying memory layer that records recent execution signals on warehouse nodes and directed edges to rank feasible endpoints and route preferences; it does this without changing collision constraints or planner validity.
  39. 62 score ArXiv · Worth reading · 1 min HoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven Reasoning Defines two quantitative metrics—Bias Intensity (BI) and Bias Harmfulness (BH)—and releases the LandmarkBias-3K benchmark (3,000 images) to measure the effect of landmark-induced bias on vision-language geo-localization models.
  40. 62 score ArXiv · Worth reading · 1 min SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions SciDiagramEdit builds a benchmark that mines before/after figure pairs from arXiv version histories and operates on the figure's editable vector source so users can inspect and co-edit individual primitives.
  41. 62 score Twitter/X · Worth reading · 1 min @alexocheema says he cannot share the original Steve Jobs Theatre slides, so he… @alexocheema says he cannot share the original Steve Jobs Theatre slides, so he described one slide and used AI to generate a recreation.
  42. 62 score Twitter/X · Worth reading · 1 min The idea for local.ai originated at Apple Park and, per @alexocheema on… The idea for local.ai originated at Apple Park and, per @alexocheema on 2026-07-16 21:46:03+00:00, was built out at NVIDIA HQ using NVIDIA-provided hardware with 'no strings attached.'
  43. 62 score Twitter/X · Worth reading · 4 min Andrew Ng (post dated 2026-06-30) frames 'loop engineering' — a buzzphrase… Andrew Ng (post dated 2026-06-30) frames 'loop engineering' — a buzzphrase amplified after viral mentions by Boris Cherny (Claude Code) and Peter Steinberger (OpenClaw) — as central to building 0-to-1 products.
  44. 62 score ArXiv · Worth reading · 1 min Quantifying Training Membership Information in the Hyperspherical Embedding Geometry of Face Recognition Models Study used a factorial design over IResNet backbone size, loss head, training duration, and number of training identities to train 180 face‑recognition models and quantify cluster-geometry membership signals.
  45. 62 score Twitter/X · Worth reading · 1 min Tori Shivanandan, Radiant President & COO, says Radiant is building microreactors… Tori Shivanandan, Radiant President & COO, says Radiant is building microreactors and that 'small details' are critical to proving the product and de‑risking it for customers.
  46. 62 score Twitter/X · Worth reading · 1 min Andrew Ng announced on 2026-05-14 a new course titled 'Transformers in Practice'… Andrew Ng announced on 2026-05-14 a new course titled 'Transformers in Practice', built in partnership with AMD and taught by Sharon Zhou.
  47. 62 score Twitter/X · Worth reading · 1 min Rebuilt a FedEx-facing delivery orchestration platform in 3.5 months with 2… Rebuilt a FedEx-facing delivery orchestration platform in 3.5 months with 2 engineers, producing 122 merged pull requests in the first 90 days vs. the client's 7–8 month rebuild estimate.
  48. 62 score ArXiv · Worth reading · 1 min AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning AlphaWiSE is a post-hoc weight-space interpolation method that composes two frozen source checkpoints by fitting one scalar interpolation coefficient per aligned parameter tensor (the scalar is shared across all entries of that tensor).
  49. 62 score Twitter/X · Worth reading · 2 min Karpathy presented three concrete 'new horizons' for LLMs at Sequoia Ascent 2026… Karpathy presented three concrete 'new horizons' for LLMs at Sequoia Ascent 2026 (talk published 2026-04-30): (1) menugen — an app wholly driven by LLMs that takes an image and outputs an image with no classical code required; (2) installing '.md' skills instead of '.sh' scripts — writing installations in English for an LLM to interpret, target, and debug for a user's setup; (3) LLM knowledge bases that enable computation over unstructured data from arbitrary sources (text/articles), a capability he says was impossible with classical code.
  50. 60 score Twitter/X · Worth reading · 1 min @swyx claims GPT-5.6 + Superapp now outperforms prior CUA efforts he tracked… @swyx claims GPT-5.6 + Superapp now outperforms prior CUA efforts he tracked: World of Bits (Shi et al. 2017), Adept/@jluan (interview ~3 years ago), Anthropic's Computer Use launch (~2 years ago), Claude Cowork (~3 months ago), and a full CUA track at @aidotengineer (3 weeks ago).
  51. 60 score ArXiv · Worth reading · 1 min When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space Hidden-state analyses across Qwen2.5-3B/7B/14B/32B, Phi-3.5, and SmolLM2 show content danger (CD) and physical danger (PD) form separable signals; PRISM (single-layer L2-regularized logistic probe over full hidden states) achieves 86.2–87.7% accuracy with 11.7–13.7% FPR on SafeAgentBench, versus 24.7–39.0% FPR for same-scale LLM judges.
  52. 60 score Twitter/X · Worth reading · 1 min Runta raised a $20M seed round led by a16z, announced on 2026-07-16. Runta raised a $20M seed round led by a16z, announced on 2026-07-16.
  53. 58 score Twitter/X · Worth reading · 1 min @emollick (2026-07-17) flags Kimi K3 as having no model card and an expected… @emollick (2026-07-17) flags Kimi K3 as having no model card and an expected weight release “in a couple weeks,” asking how pre-clearance would apply to open-weight models that are easy to jailbreak.
  54. 48 score Twitter/X · Quick skim · 1 min Richard S. Richard S. Sutton announced on 2026-07-13 that he and Khurram Javed have broken away from John Carmack (@ID_AA_Carmack) and Keen Technologies to found a new startup, Oak Lab (@oaklab_ai).
  55. 48 score Twitter/X · Quick skim · 2 min Adam Thierer warned on 2026-07-09 that the U.S. currently lacks clear statutory… Adam Thierer warned on 2026-07-09 that the U.S. currently lacks clear statutory AI frameworks, leaving an opaque, sporadic de facto model-review/licensing regime in which 'voluntary' agreements and national-security pressure can force large model developers to comply or face market removals, long delays, or lost government contracts.
  56. 48 score Twitter/X · Quick skim · 1 min The ruptured pipe was a 100-year-old, 36-inch trunk line under Sunset Boulevard… The ruptured pipe was a 100-year-old, 36-inch trunk line under Sunset Boulevard that LADWP had scheduled for replacement in 2031; LADWP budgeted $280 million to replace 6.4 miles of this exact trunk line, with design ongoing and construction planned 2031-2035—the pipe burst five years before crews would start.
  57. 42 score Twitter/X · Quick skim · 1 min Aakash Gupta (tweeted 2026-07-16) asserts AI has collapsed the defenders' time… Aakash Gupta (tweeted 2026-07-16) asserts AI has collapsed the defenders' time lag: attacker research that used to take days now takes minutes, so speed of scanning is no longer a sustainable edge—"whoever ships clean wins."
  58. 42 score ArXiv · Quick skim · 1 min Goal-Oriented Semantic Communication for Distributed ISAC-Enabled Vehicle Coordination Introduces a goal-oriented semantic communication (GSC) framework for distributed ISAC-enabled vehicle coordination at unsignalized intersections; GSC transmits sensing and C&C signals only when semantically important for improving intersection throughput.
  59. 42 score ArXiv · Quick skim · 1 min Catch, Throw, Repeat: Planning for Human-Robot Partner Juggling Presents a real-time planning and control architecture that combines predictive ball tracking, adaptive online trajectory optimization using a multiple-shooting formulation, and a state-machine coordination logic to enable synchronized multi-ball human-robot partner juggling.
  60. 42 score ArXiv · Quick skim · 1 min ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors ARMOR++ (2026-07-16) achieves a statistically confirmed, substantial gain in blind-target Attack Success Rate (ASR) over the state-of-the-art agentic baseline on the AADD-2025 benchmark, and also outperforms non-agentic baselines and defended detectors.
  61. 42 score ArXiv · Quick skim · 1 min Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation Proposes a self-evolving, expert-in-the-loop annotation framework for Major Depressive Disorder (MDD) that pairs LLM-assisted labeling with expert verification and operates in three stages: candidate evidence selection, DSM-5-TR criterion-level analysis, and case-level synthesis; outputs labels plus clinical evidence, reasoning traces, and edit histories.
  62. 40 score Twitter/X · Quick skim · 1 min Andrew Ng posted on 2026-05-20 a new deeplearning.ai short course (partnered with… Andrew Ng posted on 2026-05-20 a new deeplearning.ai short course (partnered with Google Cloud) taught by Katie Nguyen and Wafae Bakkali on building AI agents that generate images and videos, emphasizing agents that evaluate and iterate on their own outputs.
  63. 39 score Twitter/X · Quick skim · 1 min @levie (published 2026-07-15) argues code is especially amenable to agents… @levie (published 2026-07-15) argues code is especially amenable to agents because it can be quickly tested—either by manual checks or by running automated tests—giving fast feedback loops that most other work lacks.
  64. 39 score Twitter/X · Quick skim · 1 min Detroit recorded an AQI of 724 on 2026-07-17; the 'Hazardous' category starts at… Detroit recorded an AQI of 724 on 2026-07-17; the 'Hazardous' category starts at 301 and Delhi often registers around 400, and Berkeley Earth’s conversion (22 µg/m³ PM2.5 ≈ one cigarette/day) implies Detroit’s peak PM2.5 exposure equated to roughly a pack of cigarettes for a day spent outdoors.
  65. 38 score ArXiv · Quick skim · 1 min Mutable Low-Rank Sketches for Retrain-Free Recommendation Mutable sketches store each user in a KP-tree (a sparse segment tree with sum aggregation), fit one low-rank projection, and recompute embeddings on-the-fly; Theorem 1 shows each new observation monotonically tightens the prediction-error envelope (a guarantee FunkSVD and eALS lack).
  66. 35 score ArXiv · Quick skim · 1 min Decoding Market Emotion from Blockchain Activity: A Data-Driven Sentiment Classifier XGBoost was the best-performing classifier in the study, achieving an average F1-score of about 0.84 for daily Twitter-based Bitcoin sentiment labels using cross-validation (paper published to arXiv 2026-07-16).
  67. 35 score Twitter/X · Quick skim · 1 min Kimi-K3 scored 1679 points to reach #1 in Arena.ai's Frontend Code Arena… Kimi-K3 scored 1679 points to reach #1 in Arena.ai's Frontend Code Arena (announcement posted July 17, 2026), overtaking Claude Fable 5.
  68. 35 score Twitter/X · Quick skim · 1 min On 2026-07-17 Guillermo Rauch announced Pete Hunt (@floydophone) has joined… On 2026-07-17 Guillermo Rauch announced Pete Hunt (@floydophone) has joined Vercel to run Frameworks and lead Next.js; Hunt is described as an early React pioneer at Meta who powered Instagram Web adoption.
  69. 35 score Twitter/X · Quick skim · 1 min An 11-year-old girl in rural Tennessee was reported missing; a nearby Flock… An 11-year-old girl in rural Tennessee was reported missing; a nearby Flock street camera captured a car whose license plate matched a registered sex offender. Police used the plate and the vehicle's direction of travel on I‑75 to locate and rescue the bound but alive girl.
  70. 35 score ArXiv · Quick skim · 1 min ESAR: Event-Based Synthetic Aperture Reconstruction ESAR formulates monocular event-camera reconstruction as a synthetic-aperture inverse problem that recovers a static ground-domain log–radiance field θ ∈ ℝ^{N_g}, replacing the latent pixel-time volume v ∈ ℝ^{N_pN_t} with the geometric relation v = Pθ and the linearized measurement model APθ = b + η (A = temporal differencing, b = signed binned event counts).
  71. 35 score ArXiv · Quick skim · 1 min SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment SUFLECA scales geometry-grounded feature learning with Normalized Object Coordinates (NOCs) supervision on 674,000 images spanning 12 real and synthetic datasets, producing compact geometry-aware features that generalize across domains.
  72. 35 score ArXiv · Quick skim · 1 min Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA A retrospective GI endoscopy case study on MediaEval Medico 2025 compared nine documented multimodal VQA systems and found parameter-efficient adaptation of pretrained backbones produced the strongest challenge performance, but improvements in answer-level metrics did not consistently reflect faithful or complete clinical reasoning.
  73. 35 score ArXiv · Quick skim · 1 min Motion-Conditioned Multi-View Fusion for Myocardial Infarction Localization from Echocardiography MCF-Net achieves 72.4% F1 and 84.9% accuracy for segment-level myocardial infarction (MI) localization, outperforming motion-only, vision-only, and prior fusion baselines on echocardiography.
  74. 35 score ArXiv · Quick skim · 1 min QuReC: All-in-One Image Restoration with Query-Specific Guidance and Local-Global Response Calibration QuReC (Zhou et al., accepted at ACM MM 2026; arXiv:2607.15097v1, published 2026-07-16) introduces two modules: a Degradation-Guided Query Reconstruction Module (DQRM) that matches each spatial query to a degradation prototype space to produce query-specific degradation-aware representations, and a Local-Global Response Calibration Module (LGRCM) for dual-branch aggregation calibrated by learnable priors.
  75. 35 score ArXiv · Quick skim · 1 min DriftWorld: Fast World Modeling through Drifting DriftWorld learns an action-conditioned drift to generate multi-step future frames in a single forward pass, running at 30+ fps and producing rollouts on average 17× faster than diffusion-based world-model baselines.
  76. 35 score Twitter/X · Quick skim · 2 min 30 engineering teams entered the company's internal hackathon; a PM (Jyothi… 30 engineering teams entered the company's internal hackathon; a PM (Jyothi Nookula) won by implementing an adversarial-agent approach inspired by a public Anthropic blog post and shipped it into production (post published 2026-07-16).
  77. 35 score ArXiv · Quick skim · 1 min Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence Rubrics on Trial is a query-only framework that evolves a rubric set from an empty seed using only synthetic rubric-conditioned response pairs (no human-written rubrics, preference data, or model training); it validates each proposed rubric and filters out non-discriminative, over-specific, and style-only candidates.
  78. 35 score Twitter/X · Quick skim · 1 min Simon Dedic (@sjdedic) posted on 2026-07-16 that his profile-pic change is a… Simon Dedic (@sjdedic) posted on 2026-07-16 that his profile-pic change is a strategic signal marking the end of an era: after starting his crypto career in 2017 he intends to move away from pseudo-anonymous/comic pfps and embrace doxxing/accountability at Moonrock Capital.
  79. 35 score Twitter/X · Quick skim · 1 min On 2026-07-14, @fchollet praised Harvinder and Suman’s Airtap for turning SMS… On 2026-07-14, @fchollet praised Harvinder and Suman’s Airtap for turning SMS into a headless, agentic execution layer that operates mobile apps (explicitly naming DoorDash and TikTok) in the background while providing plain-text updates and only prompting the user for authentication.
  80. 35 score ArXiv · Quick skim · 1 min Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies The authors compare Grokipedia (released late 2025) and Wikipedia using 1,394 article pairs about government members, scored along nine expert-coded ideological dimensions and judged by four LLMs: Grok, Claude, Mistral, and DeepSeek (paper posted 2026-07-16).
  81. 35 score Twitter/X · Quick skim · 1 min Argentina removed 13 zeros from its currency between 1970 and 1992; one modern… Argentina removed 13 zeros from its currency between 1970 and 1992; one modern peso equals 10 trillion of the original pesos.
  82. 35 score Garry's List · Quick skim · 13 min The Merchant and the Lawyer Sinclair Louie (1950s San Francisco) bought a Sea Cliff home with legal help from Jewish lawyer Ben Lehr, a symbolic case of Chinese–Jewish solidarity that prefigured multiethnic Civil Rights coalition work.
  83. 35 score ArXiv · Quick skim · 1 min MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos MAGiSt3R is a multi-agent 3D reconstruction framework that processes monocular RGB videos at almost 10 FPS using a feed-forward 3R-family model to regress local point maps and a dedicated merging module (MAGMA) to fuse maps intra-agent and inter-agent.
  84. 35 score ArXiv · Quick skim · 1 min Ray-based phase error correction for miniaturized DOE projector-based FPP under single-directional hyperbolic projection Proposes a ray-based phase-error correction framework for miniaturized DOE projector-based Fringe Projection Profilometry (FPP) that models phase artifacts along projection rays from the projector pinhole, avoiding image-domain or neighboring-pixel processing.
  85. 35 score ArXiv · Quick skim · 1 min Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Authors Paul Kassianik, Blaine Nelson, and Yaron Singer (arXiv 2026-07-16) propose a cost-aware evaluation that compares security agents at fixed cost levels and decomposes performance into inference spend versus tool (telemetry/enrichment) spend; results and an interactive site are at https://evals.frontier.security.
  86. 35 score Twitter/X · Quick skim · 1 min Kimi K3 (an open Chinese model) beat Opus 4.8 on automated retro-game generation… Kimi K3 (an open Chinese model) beat Opus 4.8 on automated retro-game generation for Road Fighter, Battle City, and Q*bert, producing the best Q*bert where the player jumps on cubes, paints them, and flees a purple snake; Kimi reportedly got gameplay, physics, mechanics, and AI behavior all working together.
  87. 35 score Twitter/X · Quick skim · 1 min Ernesto Lopez (@ErnestoSOFTWARE) says his bootstrapped, two‑person app makes… Ernesto Lopez (@ErnestoSOFTWARE) says his bootstrapped, two‑person app makes $50,000/week across iOS and Android — roughly $200,000/month — and ranks #7 against well‑funded competitors (post dated 2026-07-15).
  88. 35 score Twitter/X · Quick skim · 1 min On 2026-07-15 Aaron Levie stated "Code is OP," arguing code's testability (you… On 2026-07-15 Aaron Levie stated "Code is OP," arguing code's testability (you can run tests or manually verify behavior) makes it uniquely amenable to agent-driven automation.
  89. 35 score ArXiv · Quick skim · 1 min DAPGNet: Dynamic Adaptive Physics-Guided Graph Diffusion Network for Hyperspectral Image Classification DAPGNet achieves state-of-the-art OA, AA, and Kappa on Indian Pines, WHU-Hi-LongKou, Houston2013, and Houston2018; it improves average accuracy (AA) by 3.64–7.31 percentage points over the strongest competing method.
  90. 35 score ArXiv · Quick skim · 1 min Concept-Guided Spatial Regularization for World Models in Atari Pong Reproduced five visual world-model agents (DreamerV3, DIAMOND, TWISTER, Simulus, STORM) on Atari Pong, froze each learned world model, and evaluated them with closed-loop rollouts driven by separately trained policies.
ArXiv · 1 min Signal

T^2MLR: Transformer with Temporal Middle-Layer Recurrence

Paper (Ziyang Cai et al., published on arXiv 2026-07-16) introduces T^2MLR (Transformer with Temporal Middle-Layer Recurrence) that fuses a cached middle-layer representation from the previous token into an earlier layer of the current token to preserve intermediate reasoning states.

T^2MLR (Transformer with Temporal Middle-Layer Recurrence) addresses autoregressive decoding’s loss of intermediate hidden computation by caching a middle-layer representation from the previous token and injecting it into an earlier layer of the current token. With little inference overhead, targeted middle-layer recurrence (sometimes only 20% of layers) outperforms matched baselines across pretraining and multi-hop reasoning; retrofitting a 1.7B pretrained model improves math reasoning.

T^2MLR consistently outperforms data- and parameter-matched Transformer baselines across natural-language pretraining and multi-hop reasoning finetuning; applying recurrence to only a localized middle-layer block — as little as 20% of the network — often beats full-layer recurrence.
Open reader
Must read

Start here. These are the items with the strongest reader value today.

1 items
Worth reading

Useful context and follow-up reading when you have more time.

52 items
1 ArXiv 2026-07-16 1 min read
Open

Pretraining Data Can Be Poisoned through Computational Propaganda

Why it matters

Graf et al. (Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith, David Kohlbrenner, Kyle Lo; arXiv 2026-07-16) show that pretraining-data poisoning is feasible via public web discussion interfaces, extending beyond prior Wikipedia‑focused attacks and identifying third‑party webpage content as a viable attack vector.

  • They introduce HalfLife, a novel analysis method to estimate whether adversarial content injected into the web is included after web crawling and data curation, and they study how poisoned injections interact with web‑crawl based LM training pipelines (abstract only).

Pretraining Data Can Be Poisoned through Computational Propaganda (Graf et al., arXiv 2026-07-16) demonstrates that attackers can introduce harmful LM behaviors by injecting content into public web discussion interfaces rather than only trusted sources like Wikipedia. The authors introduce HalfLife, an analysis for estimating whether injected content survives web crawling and curation, and show third‑party webpages are a realistic poisoning vector (based on the abstract).

Authors: Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith...
2 ArXiv 2026-07-16 1 min read
Open

BadWAM: When World-Action Models Dream Right but Act Wrong

Why it matters

BadWAM defines World-Action Drift Attacks against World-Action Models (WAMs), characterizing the attack surface along two criteria—attack strength and stealthiness—and instantiates two attacks: an action-only attack and an imagination-preserving attack.

  • Under closed-loop execution the action-only attack drops task success from 96.5% to 43.1% on evaluated WAM variants, demonstrating large end-to-end degradation.
  • The imagination-preserving attack induces harmful action shifts while keeping predicted futures close to clean imaginations; the authors show moderate future-preserving regularization can retain strong attack performance while reducing imagined-future drift, revealing a WAM-specific vulnerability.

BadWAM introduces a unified framework for World-Action Drift Attacks that break the alignment between a WAM's imagined future and its executed actions. The authors formalize attack strength vs. stealthiness and instantiate action-only and imagination-preserving attacks, evaluate them across WAM variants, report a drop from 96.5% to 43.1% success for the action-only attack, and highlight a vulnerability that moderate future-preserving regularization can partially mask.

Authors: Qi Li, Xingyi Yang, Xinchao Wang
3 ArXiv 2026-07-16 1 min read
Open

Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search

Why it matters

Using a ReAct agent over HotpotQA, the authors replayed 1,000 development questions and computed a Counterfactual Trajectory Utility (CTU) by deleting each read document and re-running the trajectory; across 23,322 document observations CTU and Static RAG Utility (SRU) are nearly independent (Spearman rho = -0.026).

  • Roughly one-third of documents the agent reads are 'bridge documents'—causally load-bearing despite appearing useless to a static reader; a BM25 + cross-encoder proxy yields a bridge cell of 27.2% on an evenly spread axis.
  • Observable Entity Relevance (OER) analysis shows discriminative entities from relevant documents appear in the agent's next query 4.02× more often than entities only in non-relevant documents (6.1% vs 1.5%, n = 227,139), indicating bridges redirect the search by supplying discriminative entities.

Agentic multi-step retrieval with a ReAct agent on HotpotQA finds static retrieval scores poorly predict a document's causal impact. The authors compute Counterfactual Trajectory Utility (CTU) by deleting each read document over 1,000 dev questions and measuring final-answer, next-query, and turn-count deltas; across 23,322 observations SRU and CTU are nearly independent (Spearman ρ = −0.026). About one-third of documents are bridge documents; OER shows discriminative entities occur 4.02× more in subsequent queries (6.1% vs 1.5%, n=227,139), implying static relevance does not equal causal usefulness.

Authors: Debayan Mukhopadhyay, Utshab Kumar Ghosh, Shubham Chatterjee
4 ArXiv 2026-07-16 1 min read
Open

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

Why it matters

LLMs often violate the law of total probability under test: using binary-tree partitions (prompting models with verbalized subpopulation descriptions and aggregating estimates) yields inconsistent population-level marginals across partitions for state-of-the-art 'frontier' models and multiple problem domains.

  • Macro fallacy discovered: population estimates reconstructed from more fine-grained persona/subpopulation prompts frequently align better with human reference data than direct population-level estimates; this effect is robust across tree structures and estimation tasks and can be partially recovered via implicit prompting.
  • The authors propose statistical self-consistency (partition-aggregate consistency) as a reference-free evaluation criterion, arguing models possess relevant subpopulation knowledge but do not reliably propagate it into aggregate estimates.

Partition, Prompt, Aggregate evaluates whether in‑context LLM outputs behave like conditional probability estimates by testing the law of total probability via recursive binary-tree partitions: prompt models on verbalized subpopulations, aggregate back to the population, and compare across partition granularities. Across problem domains and frontier models the authors find widespread inconsistencies and identify a "macro fallacy" where fine-grained (persona) prompts yield aggregates closer to human references than direct population prompts. They propose statistical self-consistency as a reference-free metric, showing models encode subpopulation knowledge but fail to propagate it reliably.

Authors: Patrik Wolf, Thomas Kleine Buening, Andreas Krause...
5 ArXiv 2026-07-16 1 min read
Open

Sharp Stability Threshold and Certification for Designing Stable Residual Architectures

Why it matters

Proposes the sublinear-growth principle and proves a sharp stability threshold q = 1 for residual-block velocity fields obeying ‖v(x,t)‖ ≤ c‖x‖^q + b (q ∈ [0,1]); classical ODE theory gives global forward flow for q ≤ 1 and divergent velocity fields for any q > 1.

  • An optimal-control / HJB analysis shows the training optimum is bang–bang on the admissible-class boundary: optima with q > 1 blow up while q ≤ 1 are safe, giving a necessary and sufficient condition for stable training and enabling an arithmetic of input-magnitude exponents across five architectural operations to certify stability at the primitive level.
  • Provides a parameter-free modification that reduces a supercritical Mamba block from q = 5 to q = 1 without layer normalization; experiments on Mamba and PatchTST confirm that q ≤ 1 variants train stably, indicating stability depends on the input-magnitude exponent rather than the mere presence of normalization.

The paper introduces the sublinear-growth principle for deep residual architectures, characterizing block velocities by ‖v(x,t)‖ ≤ c‖x‖^q + b and proving q ≤ 1 is the sharp stability threshold via ODE existence results and an HJB optimal-control selection argument. It supplies an algebra of input-magnitude exponents for five primitive operations to certify architectural stability, gives a parameter-free fix that lowers Mamba's q from 5 to 1, and reports experiments on Mamba and PatchTST showing q ≤ 1 yields stable training. Summary based on the abstract (full text not available here).

Authors: Hyemin Gu, Michael Tyrrell, Tuhin Sahai...
6 Twitter/X 2026-06-04 1 min read
Open

On 2026-06-04 Andrew Ng announced a short deeplearning.ai course (built with Red…

Why it matters

On 2026-06-04 Andrew Ng announced a short deeplearning.ai course (built with Red Hat) taught by Cedric Clyburn on serving LLMs to many concurrent users with low latency and reasonable cost.

  • A 70B-parameter model requires ~140 GB just to load weights, and every active request also needs GPU memory for a KV cache to store token context; the course teaches reducing memory footprint via quantization and serving with vLLM.
  • Practical skills taught include quantizing models and measuring accuracy tradeoffs, deploying with vLLM to handle concurrent requests, and benchmarking to balance speed, cost, and accuracy (signup: deeplearning.ai/courses/fast…).

Andrew Ng's new deeplearning.ai short course (announced 2026-06-04), produced with Red Hat and taught by Cedric Clyburn, teaches practical methods for serving LLMs at scale. It focuses on memory management—e.g., a 70B model needs ~140 GB for weights plus per-request KV caches—covering quantization, vLLM deployment, and benchmarking tradeoffs between speed, cost, and accuracy.

By @AndrewYNg
7 ArXiv 2026-07-16 1 min read
Open

In-Place Tokenizer Expansion for Pre-trained LLMs

Why it matters

Introduces in-place tokenizer expansion: continue the source tokenizer's BPE merges on a multilingual corpus, copy carried-over embedding rows unchanged, initialize new token embeddings as the mean of their source sub-token embeddings, and use a two-stage adaptation (embedding-only training, then full-model continued pre-training) to recover source-checkpoint quality.

  • Applied to a continued checkpoint of LFM2-8B-A1B to produce LFM2.5-8B-A1B with a 128K tokenizer; the expanded tokenizer encodes Hindi and Vietnamese in ~2.4× and ~2.6× fewer tokens (up to 4.0× on Thai) and yields an estimated 2.2–3.7× per-character decode speedup; model weights and the expanded tokenizer are released.

In-Place Tokenizer Expansion upgrades a pre-trained LLM's tokenizer by continuing its BPE merges on a multilingual corpus: carried-over tokens keep their embeddings, new tokens are initialized as the mean of their source sub-token embeddings, and a two-stage adaptation (embedding-only then full-model continued pre-training) restores checkpoint quality. Applied to LFM2-8B-A1B → LFM2.5-8B-A1B with a 128K tokenizer, Hindi/Vietnamese/Thai see 2.4×/2.6×/up to 4.0× token reductions and estimated 2.2–3.7× per-character decode speedups; weights and tokenizer are released.

Authors: Jimmy T. H. Smith, Tarek Dakhran, Alberto Cabrera...
8 Twitter/X 2026-05-31 8 min read
Open

On 2026-05-31 Richard S.

Why it matters

On 2026-05-31 Richard S. Sutton asserts that generative AI trained by supervised learning (LLMs, image/video models, world models) cannot make novel discoveries because it lacks runtime Evaluation and thus cannot perform selective retention — stochastic generation yields novelty but not evaluated, retained discoveries.

  • Sutton defines 'Discovery' as the combination of three steps — Variation, Evaluation, and Selective retention — and maps these to known mechanisms (reinforcement learning, instrumental learning/operant conditioning, generate-and-test) rather than to plain supervised/backprop learning.
  • He cites concrete examples of systems that achieved both novelty and quality by having explicit evaluation/objectives: AlphaGo (move 37), AlphaZero, GT-Sophy (simulated racecar), AlphaFold, AlphaProof, Claude-Code, and RL-Lyft — all leverage search, RL, or objective-driven evaluation.
  • Sutton notes a weakness of standard backprop (random initialization supplies only one-time variation) and highlights his group's 'continual backpropagation' (Nature, described as published a couple of years prior) that periodically re-initializes underused neurons to preserve variation and plasticity; he calls for sharing explicit goals with AIs to enable automated creativity and discovery.

Richard S. Sutton (video/text published 2026-05-31) argues that mainstream generative AI—large language, image, and video models trained by supervised learning—can produce outputs that are novel or good but not truly creative discoveries because they lack a runtime Evaluation and selective retention loop. He formalizes Discovery as Variation + Evaluation + Selective retention, equating Evaluation with objectives or reward signals found in reinforcement learning, search, and the scientific method. Sutton contrasts mimicry-based generative models with systems that have produced real advances (AlphaGo’s move 37, AlphaZero, GT-Sophy, AlphaFold, AlphaProof, Claude-Code, RL-Lyft), which use explicit evaluation to keep novel, high-value results. He further argues that backprop’s randomness is typically limited to initialization and highlights his group’s “continual backpropagation” (Nature) that reinitializes underused neurons to maintain plasticity, concluding with a call to grant AIs explicit goals so they can autonomously create and discover.

By @RichardSSutton
9 Twitter Article 2026-04-11 8 min read
Open

T0nyav’s 11 April 2026 post frames Anthropic’s Mythos and Project Glasswing as…

T0nyav’s 11 April 2026 post frames Anthropic’s Mythos and Project Glasswing as evidence the internet’s permissionless frontier is closing: Mythos won’t be generally released and access appears concentrated among enterprise partners (AWS, Apple, Cisco, CrowdStrike, Google, Microsoft, NVIDIA, etc.). Citing Frederick Jackson Turner (1893), Rudolf Laine (2024) and George Hotz (2026), the author warns that privatizing frontier intelligence creates permanent advantages, a “zero day generator,” and state-scale capabilities without public accountability. They argue broader API access and open-source models improve safety by exposing latent capabilities for real-world testing (models are “eval-aware”), noting MATS symposium data where roughly two-thirds of posters used Chinese open-source models and that open models currently lag frontier systems by about 3–12 months. The piece calls for government-like governance of labs: public access criteria, due process, appeals, and FOIA-style audits, while acknowledging this might be a temporary mainframe-era phase before cheap local models proliferate.

By t0nyav
10 Twitter Article 2026-04-02 1 min read
Open

Mintlify's AI assistant now uses ChromaFS, a virtual filesystem that turns UNIX…

Mintlify's AI assistant now uses ChromaFS, a virtual filesystem that turns UNIX commands into database queries to simulate a sandbox. Announced 2026-04-02 by @cdxker, this replaced chunk-based RAG (which missed cross-page context) and avoided slow sandboxes (~46s startup, $70k/yr), cutting latency to ~10ms and serving 30,000+ conversations daily.

By cdxker
11 ArXiv 2026-07-16 1 min read
Open

Data Driven Block Replacement Scheduling

Why it matters

Presents Hoeffding- and Bernstein-based lower-confidence-bound bandit algorithms for choosing replacement interval k ∈ {1,…,K}, achieving O(K log T) regret and matching the Lai–Robbins lower bound; correlated variants attain O((K−k*) log T) regret and require only O(1) direct pulls of suboptimal arms k < k*.

  • Proposes a Kaplan–Meier renewal estimator that learns the lifetime distribution from censored data, proving almost-sure policy consistency and reporting empirically near-zero incremental regret at long horizons; two average-cost MDP analyses show block replacement is optimal within a time-elapsed policy class and give a monotone threshold age-vector benchmark under increasing failure rates.

The paper studies learning the cost-minimizing block-replacement interval k* for N identical machines when lifetimes are unknown, modeling choices k ∈ {1..K} as a stochastic multi-armed bandit with censored lifetime observations. It gives LCB algorithms (Hoeffding/Bernstein) with O(K log T) regret (Lai–Robbins optimal), correlated-arm refinements with O((K−k*) log T), a Kaplan–Meier renewal estimator with almost-sure consistency, and MDP analyses proving block replacement optimality and a threshold structure under increasing failure rates.

Authors: Aniruddhan Ganesaraman, VIdyadhar Kulkarni
12 ArXiv 2026-07-16 1 min read
Open

What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity

Why it matters

The paper proves the 2025 conjecture of Patel et al. that a bounded second-order heterogeneity assumption yields improved convergence guarantees for Local SGD (Federated Averaging) on general convex objectives; the authors (K. K. Patel, R. Islamov, S. U. Stich, A. Lucchi, E. Gorbunov, L. Wang) give improved upper bounds and nearly-tight lower bounds (arXiv:2607.14731v1, published 2026-07-16).

  • As an additional contribution, the authors derive a new lower bound for serial (with-replacement) SGD showing that second-order heterogeneity quantitatively captures the impact of rare high-curvature clients, clarifying when and why local updates can outperform minibatch SGD.

Local SGD (Federated Averaging) under a bounded second-order heterogeneity model: the authors extend prior strong-convex results to general convex objectives, proving improved convergence guarantees and producing nearly-tight matching lower bounds. Techniques also yield a lower bound for serial SGD with replacement that links rare high-curvature clients to degraded rates. Summary is based on the abstract; full text was not available.

Authors: Kumar Kshitij Patel, Rustem Islamov, Sebastian U Stich...
13 ArXiv 2026-07-16 1 min read
Open

Subjective Risk Decomposition: A New View for Uncertainty Quantification

Why it matters

Subjective-risk decomposition: Alamri, Caprio, and Brown (arXiv:2607.15196v1, published 2026-07-16) show epistemic and aleatoric uncertainty can be derived as consequences of modelling choices by decomposing a subjective risk specified by a strictly proper loss; reverse cross-entropy recovers classic information-theoretic uncertainty terms.

  • Practical and theoretical impact: the framework subsumes numerous existing UQ measures, prescribes that given a modelling scenario + strictly proper loss the epistemic/aleatoric terms are induced, and introduces subjective-risk analogues of excess risk, approximation error, and estimation error (27-page paper).

Subjective Risk Decomposition presents a framework that derives epistemic and aleatoric uncertainty as consequences of modelling choices by decomposing a subjective risk defined via a strictly proper loss. Using reverse cross-entropy it recovers classic information-theoretic uncertainty terms and subsumes many prior UQ measures; it also proposes learning-theoretic analogues (excess, approximation, estimation error).

Authors: Raghad Alamri, Michele Caprio, Gavin Brown
14 ArXiv 2026-07-16 1 min read
Open

On-Policy Delta Distillation

Why it matters

Introduces the delta signal as a distillation reward: delta = (teacher model) − (base model before instruction tuning), designed to capture changes induced by reasoning tuning; the method is named On-Policy Delta Distillation (OPD²).

  • Empirical claim: OPD² consistently outperforms conventional on-policy distillation on mathematics, science, and code-reasoning benchmarks and enables reasoning LLMs to reach strong performance with only a short post-training period; arXiv:2607.15161v1 (2026-07-16), 19 pages, 4 figures, 12 tables, code at https://github.com/naver-ai/opd2.

On-Policy Delta Distillation (OPD²) replaces direct teacher-imitation with a delta signal — the difference between a teacher and its pre-instruction-tuned base — as an on-policy distillation reward to better transfer reasoning capabilities. The authors (Heo, Hwang, Yun, Han) report consistent gains on math, science, and code-reasoning benchmarks and emphasize rapid post-training convergence; summary based on the provided abstract (full text not reviewed).

Authors: Byeongho Heo, Jaehui Hwang, Sangdoo Yun...
15 ArXiv 2026-07-16 1 min read
Open

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

Why it matters

SearchOS reformulates open-domain information seeking as relational schema completion with grounded citations and externalizes progress into Search-Oriented Context Management (SOCM) components: Frontier Task, Evidence Graph, Coverage Map, and Failure Memory.

  • System and scheduling innovations — pipeline-parallel task scheduling, a Search Tool Middleware Harness, and a hierarchical skill system (strategy + access skills) — reduce repeated failed searches and improve utilization; SearchOS outperformed all evaluated single- and multi-agent baselines on the WideSearch and GISA benchmarks.
  • Paper (arXiv:2607.15257v1, 2026-07-16) provides code at https://github.com/antins-labs/SearchOS.

SearchOS is a system-level multi-agent framework that tackles open-domain information seeking by turning implicit search progress into explicit, shared state. The authors cast the task as relational schema completion with grounded citations and introduce SOCM (Frontier Task, Evidence Graph, Coverage Map, Failure Memory), pipeline-parallel scheduling, a Search Tool Middleware Harness, and a hierarchical skill system to avoid repetition and improve throughput. SearchOS leads all metrics versus single- and multi-agent baselines on WideSearch and GISA; code is released on GitHub.

Authors: Yuyao Zhang, Junjie Gao, Zhengxian Wu...
16 ArXiv 2026-07-16 1 min read
Open

Mask-Aware Policy Gradients for Diffusion Language Models

Why it matters

Proposes mask-aware policy gradients for Masked Diffusion Language Models (MDLMs), formalizing generation as a two-stage action MDP (token placement then which positions to remask) and showing the policy gradient decomposes into a token term and a masking term.

  • Optimizing both token and masking terms yields state-of-the-art results on reasoning and coding benchmarks: 87.1% on GSM8K and 53.4% on MBPP; paper by Haran Raajesh, Kulin Shah, Adam Klivans, and Philipp Krähenbühl was accepted at COLM 2026.
  • Identifies a shortcoming of prior MDLM RL methods that approximate the log-likelihood by modeling only token predictions and ignore the order of position unmasking; the new method explicitly models both decisions.

Mask-aware policy gradients formalize MDLM generation as a two-stage MDP—choosing tokens and choosing which positions to remask—and derive a policy-gradient decomposition into token and masking terms. Optimizing both terms produces state-of-the-art performance on math and coding benchmarks (87.1% GSM8K, 53.4% MBPP). Only the abstract was provided here; full text was not included.

Authors: Haran Raajesh, Kulin Shah, Adam Klivans...
17 ArXiv 2026-07-16 1 min read
Open

The Industrialization of Research ; On AI-Driven Science and Its Consequences

Why it matters

Frames 'industrialization of research' as a shift from a craft model to an automated, supervised pipeline; cites the US Department of Energy's Genesis Mission as the most ambitious current instantiation (ArXiv:2607.15164v1, published 2026-07-16).

  • Enumerates seven specific risks posed by AI-driven science: erosion of intergenerational transmission of scientific competence; opacity of AI-generated theories; collapse of peer evaluation from a flood of machine output; uncertain capacity for paradigm-shifting discovery; capture of agendas by political/industrial actors; compounding systematic errors in closed-loop pipelines; and structural bifurcation of the global research community.
  • Author Emmanuel Jeannot does not oppose AI-driven science but presents these seven concerns as conditions under which AI's demonstrated potential (and risks) should be pursued responsibly.

Emmanuel Jeannot's 2026 essay 'The Industrialization of Research' argues that AI is transforming science from a researcher-centered craft into an automated, supervised pipeline, highlighting the US DOE's Genesis Mission as a prime example. The abstract lists seven concrete concerns—from skill transmission loss to peer-review collapse and systemic error amplification—and frames them as prerequisites for responsibly realizing AI-driven science. Full text was not available for this briefing.

Authors: Emmanuel Jeannot
18 ArXiv 2026-07-16 1 min read
Open

Delocalization of bias in unadjusted Hamiltonian Monte Carlo and underdamped Langevin

Why it matters

For unadjusted Hamiltonian Monte Carlo and underdamped Langevin, controlling the W2 bias of any K-dimensional marginal of a d-dimensional target requires O(√K) integration steps (up to log d factors) under assumptions of weak or sparse interactions.

  • The paper extends delocalization of bias from overdamped Langevin to discrete-time integrators, introduces a matrix-polynomial framework to analyze propagators, and proves the underdamped result holds for all large friction parameters—implying the Leimkuhler–Matthews integrator also exhibits delocalization while avoiding Metropolis cost.

The authors show that the delocalization phenomenon (previously proved for overdamped Langevin) holds for unadjusted Hamiltonian Monte Carlo and underdamped Langevin: O(√K) integration steps (up to log d) suffice to control W2 bias of any K-dimensional marginal under weak or sparse interactions. They overcome discrete-time difficulties with a matrix-polynomial propagator framework and prove results valid for all large friction, implying the Leimkuhler–Matthews integrator shares this bias delocalization.

Authors: Yifan Chen, Xiaoou Cheng, Jonathan Niles-Weed...
19 ArXiv 2026-07-16 1 min read
Open

RTS Smoother-Guided Learning of Physics-Based Neural Differential Models

Why it matters

Introduces an RTS-smoother-guided hybrid neural–physics ODE learning scheme that alternates between (1) latent-state inference with a Rauch–Tung–Striebel (RTS) smoother treating model parameters as fixed and (2) neural-network parameter updates via backpropagation on the smoothed trajectories; iterations continue until a stopping criterion.

  • Evaluated on benchmark linear, nonlinear, and stiff dynamical systems under partial state observation; method preserves interpretable mechanistic structure while improving latent-state reconstruction and long‑horizon prediction compared to pure black‑box approaches (quantitative metrics not reported in the abstract).
  • Authored by Ahmet Demirkaya, Georgios Stratis, Tales Imbiriba, Zachary D. Danziger, and Deniz Erdogmus; posted to arXiv 2026-07-16 as arXiv:2607.15180v1 (cs.LG, eess.SY) with PDF available on arXiv.

An RTS-smoother-guided hybrid neural–physics framework learns missing ODE components by alternating latent-state inference (Rauch–Tung–Striebel smoother) and neural-parameter learning (backpropagation on smoothed trajectories). Evaluated on linear, nonlinear, and stiff benchmarks under partial observations, the method retains mechanistic structure and improves latent-state reconstruction and long-horizon prediction. Only the abstract was provided.

Authors: Ahmet Demirkaya, Georgios Stratis, Tales Imbiriba...
20 Twitter/X 2026-07-16 1 min read
Open

Author @adxtyahq (published 2026-07-16) frames the prompt “design Claude Code…

Why it matters

Author @adxtyahq (published 2026-07-16) frames the prompt “design Claude Code from scratch” as an Anthropic interview question and presents an eight-step blueprint, while noting “Probably not how Claude Code is actually built.”

  • Core technical steps: build an AST + dependency graph to extract symbols/imports and cross-file relationships; use embeddings + graph traversal to retrieve only relevant files (don’t send full codebase); plan edits by identifying which files need changes and breaking tasks into small executable steps; produce minimal diffs to preserve architecture, naming, and formatting.
  • Operational practices: validate every change with linting, type checks, and tests (failed validation triggers another reasoning pass); treat search, terminal, git, and diagnostics as callable tools; maintain session memory of past edits/decisions; and explain each edit with tool outputs and validation results to build developer trust.

Author @adxtyahq (2026-07-16) proposes an eight-step design for an AI coding agent in response to the “design Claude Code from scratch” interview prompt: use ASTs and dependency graphs for program understanding, embeddings+graph traversal for targeted retrieval, explicit planning and minimal diffs for edits, continuous validation (lint/type/tests), callable tools (search/terminal/git), session memory, and explicit explanations for each change. The author adds this may not reflect Claude Code’s real implementation.

By @adxtyahq
21 ArXiv 2026-07-16 1 min read
Open

Online Neural Space Time Memory for Dynamic Novel View Synthesis

Why it matters

Proposes Online Neural Space Time Memory (NST-Mem) that decouples memory update and application: periodic (infrequent) gradient-based memory updates plus per-frame memory application via cross-view attention to handle deformations between stored memory and current frames (Elmieh et al., arXiv 2026-07-16).

  • Introduces two mechanisms—Memory Loss to force persistent internalization and Memory Caching to regularize active weights—enabling real-time, state-of-the-art novel view synthesis on dynamic human-motion scenes and minute-scale online memorization while avoiding per-frame Test-Time Training (TTT) updates.

The paper addresses real-time online novel view synthesis from multi-view streaming video by reducing the cost and instability of per-frame, gradient-based Test-Time Training updates. It decouples update/application frequencies: apply a persistent neural space–time memory every frame using cross-view attention, and perform periodic gradient memory updates. Two contributions—an auxiliary Memory Loss and Memory Caching—prevent catastrophic drift and lock in long-horizon context. According to the abstract, this yields real-time, state-of-the-art results on dynamic human motion and minute-scale memorization (preprint: arXiv 2607.15271v1, 15 pages; project demos available).

Authors: Baback Elmieh, Lynn Tsai, Zeman Li...
22 Twitter/X 2026-07-17 2 min read
Open

Saronic announced Port Alpha

Why it matters

Saronic announced Port Alpha: a $3.2 billion greenfield shipyard in Brownsville, Texas on 835 acres (expandable to 4,400 acres), built for software-defined, robotics-enabled production and autonomy; opening in 2028 to produce 850-foot vessels initially and 1,200-foot ships later, with projected impacts of >$160 billion for Cameron County, $264.5 billion for Texas, and up to 10,000 direct jobs.

  • Strategic context: a leaked Office of Naval Intelligence slide put annual U.S. shipbuilding capacity under 100,000 tons versus China’s 23.2 million tons (a ~232x gap); the Navy Secretary told Congress in 2023 that a single Chinese shipyard out-builds the entire U.S. industry, motivating Saronic to pursue greenfield construction rather than retrofitting old yards.
  • Autonomy and rapid prototyping validation: Saronic’s Louisiana yard produced a 180-foot autonomous vessel from design to launch in under a year, the company argues removing crew infrastructure enables mass production of warships, and its sea drones saw U.S. combat use for the first time this week.

Saronic's Port Alpha project is a $3.2 billion greenfield shipyard in Brownsville, Texas—835 acres expandable to 4,400—designed for software-defined, robotic production and autonomy, opening 2028 to build 850-ft ships (later 1,200-ft). The startup cites a 232× U.S.–China capacity gap (U.S. <100k tons vs China 23.2M) and recent combat use of its sea drones.

By @aakashgupta
23 ArXiv 2026-07-16 1 min read
Open

Linear representations of grammaticality in neural language models

Why it matters

Li & Kim (2026-07-16; arXiv:2607.15175v1) apply mass-mean probing and find grammatical vs. ungrammatical sentences are linearly separable in sentence representations of a wide range of pretrained neural language models; this representational separation is not fully explained by correlated sentence-level factors (e.g., lexical frequency, plausibility, world knowledge).

  • The grammaticality signal generalizes across a broad set of grammatical phenomena and, to some degree, across languages, providing a complementary, non-probability-based framework for evaluating syntactic competence in NLMs.

Li and Kim investigate whether grammaticality is encoded in internal sentence representations of pretrained neural language models. Using mass-mean probing on multiple models, they report robust linear separation between grammatical and ungrammatical strings that cannot be fully attributed to confounds like frequency or plausibility, and that generalizes across phenomena and partly across languages. Summary is based on the paper's abstract.

Authors: Jane Li, Najoung Kim
24 ArXiv 2026-07-16 1 min read
Open

RoboTTT: Context Scaling for Robot Policies

Why it matters

RoboTTT scales visuomotor context to 8K timesteps (three orders of magnitude beyond prior policies) without increasing inference latency, enabling one-shot in-context imitation, on-the-fly policy improvement, robustness to perturbations, and better long-horizon multi-stage performance.

  • On real-robot manipulation benchmarks, RoboTTT yields an 87% overall improvement versus a single-step context baseline and fully completes a five-minute, ten-stage assembly task that no baseline completes.
  • The method integrates Test-Time Training (fast weights updated by gradient descent during training and inference) into Vision-Language-Action policies and uses sequence action forcing with truncated backpropagation through time to scale context; a model pretrained with 8K timesteps outperforms the same model pretrained with 1K by 62%.

RoboTTT introduces Test-Time-Training Robot Policies that extend visuomotor context to 8K timesteps—≈1000× prior models—by representing recurrent state as fast weights updated via gradient descent during training and inference. Using sequence action forcing and truncated BPTT, it scales without added inference latency and yields an 87% improvement on real-robot manipulation, completes a five-minute ten-stage assembly, and shows 62% gain over 1K-context pretraining.

Authors: Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng...
25 ArXiv 2026-07-16 1 min read
Open

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

Why it matters

MeanFlowNFT adapts DiffusionNFT's forward-process RL to MeanFlow average-velocity generators by deriving an induced instantaneous-velocity predictor via a MeanFlow identity, enabling reward optimization while preserving MeanFlow's fast few-step (average-velocity) sampling.

  • The paper proves MeanFlowNFT inherits DiffusionNFT's strict policy-improvement guarantee and reports consistent image/video gains, outperforming prior RL-tuned few-step generators on 6 of 8 metrics on SD3.5-M.
  • Empirical highlight: 4-step MeanFlowNFT on Wan 2.1 achieves VBench 84.33 versus 50-step LongCat-Video RL at 82.57; code and models available (GitHub/Hugging Face), arXiv:2607.15273 (posted 2026-07-16).

MeanFlowNFT extends forward-process RL to MeanFlow average-velocity generators by deriving an induced instantaneous-velocity predictor from a MeanFlow identity and applying the DiffusionNFT objective; sampling still uses average velocities for fast few-step generation. The authors prove a policy-improvement guarantee and report state-of-the-art image/video results (e.g., 4-step VBench 84.33 vs 50-step 82.57).

Authors: Yushi Huang, Xiangxin Zhou, Jun Zhang...
26 ArXiv 2026-07-16 1 min read
Open

AutoSynthesis: An agentic system for automated meta-analysis

Why it matters

AutoSynthesis is an end-to-end multi-agent system that, given a natural-language research question, automates search-strategy formulation, literature retrieval, screening, full-text eligibility assessment, quantitative-statistic extraction, standardized effect-size computation, and random-effects meta-analysis.

  • In the reported application AutoSynthesis screened over 28 studies and extracted more than 20 quantitative claims; its pooled effect estimates closely match Hedges' g from expert-conducted meta-analyses, indicating close agreement with manual evidence synthesis.
  • The system also supports heterogeneity analysis, risk-of-bias assessment, and produces transparent, PRISMA-aligned reports to improve scalability of quantitative evidence synthesis.

AutoSynthesis is a 2026 multi-agent pipeline that automates quantitative meta-analysis end-to-end: it builds search strategies, retrieves and screens studies, extracts numeric results, computes standardized effect sizes, and runs random-effects meta-analysis while supporting heterogeneity and bias assessment. In one application it screened >28 studies and extracted >20 claims; pooled estimates aligned with Hedges' g from expert meta-analyses. Summary based on the abstract (full text not reviewed).

Authors: Moein Taherinezhad, Sebastian Maier, Gerardo Vitagliano...
27 ArXiv 2026-07-16 1 min read
Open

Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents

Why it matters

The paper builds a multi-agent framework combining Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and per-party Retrieval-Augmented Generation (RAG) to create manifesto-bound partisan LLM agents; DPO gives aggressive party-specific personas while RAG enforces manifesto grounding and MILT (Multi-Layered Information Lineage Topology) traces every clause into five provenance states.

  • Applied to the 2019 Flemish election in a hub-and-spoke negotiation with a formateur, three independent simulations produced a stable ranking (N-VA ahead of CD&V and Open Vld); the authors introduce a Coalition Influence Score (CIS) and report that manifesto-anchored lineage predicts real-world materialization whereas hallucinated provisions do not.

Digital Pantheon presents a transparent multi-agent testbed for coalition formation that counters RLHF-induced neutrality by combining SFT, DPO and RAG so agents remain partisan yet fact‑grounded. The system adds MILT (five provenance states) and a Coalition Influence Score to trace and quantify party contributions; on the 2019 Flemish election (three runs) it yields stable winners and shows manifesto-anchored clauses map to real-world outcomes. (Summary based on the abstract; full text not examined.)

Authors: Dylan Van Mulders, Matthias Bogaert, Dirk Van den Poel
28 Twitter/X 2026-07-16 1 min read
Open

@adxtyahq (posted 2026-07-16 19

Why it matters

@adxtyahq (posted 2026-07-16 19:21:53+00:00) claims Kimi K3 "feels like China just open-sourced a Fable-5 class model" and reports it generated exceptionally polished UI from a single prompt.

  • Kimi.ai announces Kimi K3: 2.8 trillion parameters, 1,000,000-token context, native multimodal; Kimi Delta Attention claimed to enable up to 6.3x faster decoding in million-token contexts; Attention Residuals claim ~25% higher training efficiency at <2% extra cost. Kimi K3 is live on Kimi.com, Kimi Work, Kimi Code and the Kimi API (platform.kimi.ai); weights to be opened July 27, 2026.

Kimi K3 is presented as a 2.8T-parameter, native multimodal model with a 1,000,000-token context and new techniques—Kimi Delta Attention (up to 6.3x faster decoding in long contexts) and Attention Residuals (~25% training efficiency at <2% cost). The author (@adxtyahq) on 2026-07-16 praises its single-prompt UI polish and likens it to an open-sourced Fable-5 class model; weights open July 27, 2026.

By @adxtyahq
29 Twitter/X 2026-07-17 5 min read
Open

Kimi K3 is a 2.8 trillion-parameter native multimodal open model with 1M-token…

Why it matters

Kimi K3 is a 2.8 trillion-parameter native multimodal open model with 1M-token context, a latent MoE that activates 16 of 896 experts, KDA attention and attention residuals, quantization-aware training (mxfp4 weights, mxfp8 activations), and claimed ~2.5x scaling efficiency over K2; pricing: $0.30/mtok cache-hit, $3.00/mtok cache-miss, $15.00/mtok output, and >90% cache-hit rate on coding workloads.

  • In a three-prompt procedural three.js benchmark run on aimlapi (photorealistic European roulette, Las Vegas slot machine, full 3D pinball), cost per run was Grok 4.5 $0.30, Kimi K3 $0.71, GPT-5.6 Sol $2.05, Fable-5 $7.69; Kimi produced 157,999 tokens, 2,255 lines, and took 75.6 minutes versus Grok's 34,241 tokens, 3,047 lines, and 5.1 minutes.
  • Kimi succeeded on the roulette with high-fidelity procedural wood grain and correct European sequence, produced the best pinball artwork and explicit physics reasoning (derived a 480 Hz substep, ball-settle conditions, termination guarantees), and was the only model to anticipate the three.js importmap trap.
  • Kimi failed two of three prompts: slot rendered reels backwards and used an old three.js build that ignored transmission, pinball assembled vertically at 90° with floating legs and z-fighting; 81% of Kimi's output tokens were reasoning (vs Grok 22%), making it slow and increasing wall-clock cost; price per 100 shipped lines: Grok $0.010, Kimi $0.031, Sol $0.067, Fable $0.394.

Kimi K3, a 2.8T-parameter open multimodal model released as the first 3T-class open model, was pitted against GPT-5.6 Sol, Fable-5 and Grok 4.5 on three demanding procedural three.js generation tasks (roulette, slot machine, pinball) in a 2026 head-to-head on aimlapi. Kimi shines at deep reasoning: it derived a 480 Hz physics substep, handled ball-settle logic, produced the best pinball artwork, and avoided an importmap trap others missed. But Kimi is slow and brittle—75.6 minutes for the suite, failing the slot and pinball assembly (90° cabinet, floating legs, z-fights). Cost and token metrics favored Grok for speed and delivered lines; the test illustrates open-source models closing the gap on frontier systems but still trading reliability and throughput for richer internal reasoning.

By @adxtyahq
30 ArXiv 2026-07-16 1 min read
Open

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

Why it matters

Benchmarked six MLLMs (three closed-source, three open-source) on the Scientific Visualization Literacy Assessment: 49 items, 18 scientific visualizations, 8 techniques, 11 task types, using a closed-world protocol and a human baseline from 485 participants (arXiv preprint published 2026-07-16).

  • Gemini was the strongest model, exceeding the human mean across evaluated subsets; open-source models remained below the human baseline. Models excelled at scientific illustration, search, and spatial understanding but struggled on texture‑based and integration visualizations and on quantitative estimation (notably fine-grained quantitative estimation, flow‑direction, and grounded encoding interpretation).

Multimodal large language models (MLLMs) were evaluated for scientific visualization literacy using a standardized assessment (49 items over 18 visualizations, 8 techniques, 11 task types) and human data (n=485) under a closed-world protocol. Gemini outperformed the human mean while open-source models fell below the baseline; errors concentrate on quantitative estimation, flow-direction, and grounded encoding. Code and outputs are public.

Authors: Patrick Phuoc Do, Chau M. Ta, Chaoli Wang
31 Twitter/X 2026-07-17 1 min read
Open

Alex Cheema says Kimi K3 will require 4 × 512GB M3 Ultra Mac Studios (noting 2TB…

Why it matters

Alex Cheema says Kimi K3 will require 4 × 512GB M3 Ultra Mac Studios (noting 2TB @ 3.2TB/s) and he expects it can run at 30+ tokens/sec using MTP + tensor parallelism over RDMA via Thunderbolt 5.

  • He warns model prefill will be slow on current hardware but expects a 512GB M5 Ultra (arriving in October) to make prefill ~5× faster.
  • Cheema already runs Kimi K2.5 at 24 tokens/sec on 2 × 512GB M3 Ultra Mac Studios connected with Thunderbolt 5 (RDMA) using the exolabs/MLX backend, can run 'clawdbot', and will publish local.ai benchmarks and early-access codes when weights drop.

Alex Cheema says Kimi K3 will need four 512GB M3 Ultra Mac Studios (2TB @ 3.2TB/s) and estimates 30+ tokens/sec using MTP plus tensor parallelism over RDMA on Thunderbolt 5. He expects prefill to be slow until a 512GB M5 Ultra (due in October) yields ~5× faster prefill, and promises local.ai benchmarks and early-access codes; he already runs Kimi K2.5 at 24 tok/sec on two 512GB M3 Ultras (exolabs/MLX).

By @alexocheema
32 ArXiv 2026-07-16 1 min read
Open

AHEAD: Anticipatory Hand-Driven Teleoperation via Human Intent Prediction

Why it matters

AHEAD (real-time VR teleoperation) uses a short window of 3D hand and head signals plus scene context with an attention-based classifier and a state machine in a digital twin; Top-1 intent accuracy is 76% for grasp-object prediction and 76% for target-slot prediction.

  • In a user study, AHEAD reduced robot reaction latency by 0.6 s for object selection and 1.4 s for slot selection versus baselines and participants reported lower operator workload (paper accepted to IROS2026; authors: Seok Joon Kim et al.; arXiv:2607.15172v1).

AHEAD is a VR teleoperation system that anticipates operator intent for repetitive pick-and-place by processing short windows of 3D hand and head signals plus scene context through an attention-based classifier; a state machine turns predictions into stable robot goals. The model attains 76% top-1 accuracy for both grasp and slot prediction and lowers reaction latency by 0.6–1.4 s in a user study. Summary based on the abstract; full paper PDF is available on arXiv.

Authors: Seok Joon Kim, Junho Lee, Federica Spinola...
33 Twitter/X 2026-04-09 2 min read
Open

Karpathy identifies two distinct audiences

Why it matters

Karpathy identifies two distinct audiences: (1) people who tried the free ChatGPT tier last year and judge AI by its visible quirks/hallucinations, and (2) paying technical professionals who use state-of-the-art agentic models (he names OpenAI Codex and Anthropic/Claude Code) professionally in programming, math, and research.

  • He claims state-of-the-art paid models (he references ~$200/month tiers) can dramatically outperform expectations in technical domains — “melt” programming problems that normally take days/weeks, and can spend an hour to coherently restructure an entire codebase or find and exploit vulnerabilities.
  • Free and older models (e.g., ChatGPT free tier and OpenAI’s Advanced Voice Mode) still fumble simple queries — Karpathy cites viral failures like answering “should I drive or walk to the carwash” — which fuels a perception gap on capability.
  • Karpathy attributes the technical leap to reinforcement learning with verifiable reward signals (e.g., unit tests pass/fail) and company hill-climbing toward high B2B value features, producing intense reactions he calls “AI Psychosis” among heavy coder users.

Karpathy argues there is a widening gap in public understanding of AI capabilities driven by two divergent experiences. Posted 2026-04-09, he says casual users who tried free ChatGPT last year or watched viral reels of OpenAI’s Advanced Voice Mode stumbling at simple queries form one narrative of limited capability, while professional users paying high-tier services (he references ~$200/month and names OpenAI Codex and Claude Code) experience a very different reality: agentic models that can “melt” days‑to‑weeks programming tasks, restructure codebases in an hour, and even find/exploit vulnerabilities. He explains this split by pointing to reinforcement learning’s advantage in tasks with verifiable reward signals (unit tests yes/no) and firm incentives to prioritize lucrative B2B use cases, producing intense belief among technical users that isn’t visible to casual observers.

By @karpathy
34 ArXiv 2026-07-16 1 min read
Open

Tamed Stochastic Gradient Hamiltonian Monte Carlo

Why it matters

Wang and Zhang (arXiv:2607.14862v1, published 2026-07-16) propose tamed stochastic gradient Hamiltonian Monte Carlo (tSGHMC) for sampling and stochastic optimization when stochastic gradients grow superlinearly.

  • Under a continuity-in-average condition and strong convexity, the paper proves a non-asymptotic Wasserstein-2 error bound for tSGHMC with convergence rate 1/4 and derives an upper bound on the associated expected excess risk.
  • Empirical tests on a newsvendor problem and Conditional Value-at-Risk (CVaR) minimization using synthetic and real datasets show tSGHMC attains lower root-mean-square error and lower expected excess risk than the tamed unadjusted stochastic Langevin algorithm (first-order counterpart).

The paper introduces tSGHMC to handle sampling and optimization problems with superlinearly growing stochastic gradients, proving a non-asymptotic W2 error bound with rate 1/4 under continuity-in-average and strong convexity and providing an expected excess risk upper bound. Experiments on newsvendor and CVaR tasks (synthetic and real data) show lower RMSE and excess risk versus tamed ULA.

Authors: Zhuoran Wang, Ying Zhang
35 ArXiv 2026-07-16 1 min read
Open

Hierarchical Denoising For Multi-Step Visual Reasoning

Why it matters

HDR (Hierarchical Denoising for Visual Reasoning) uses tree-structured video latents and a sparse hierarchical attention pattern (SHAP) to enable coarse-to-fine, multi-step visual reasoning; evaluated on a level-stratified benchmark covering six tasks: maze navigation, Tower of Hanoi, one-line drawing, sliding puzzle, Sokoban, and water pouring.

  • Against streaming autoregressive diffusion baselines, HDR raises success from 34.22 to 60.29 (a 76.2% relative gain) and increases average progress from 76.00 to 89.56.
  • HDR maintains low-latency streaming at 0.70 seconds per latent (54.2× faster than bidirectional diffusion) and retains 82.9% of full-data performance using only 2% training data, compared with 52.0% retention for bidirectional diffusion.

Hierarchical Denoising for Visual Reasoning (HDR) integrates tree-structured video latents and a sparse hierarchical attention pattern (SHAP) into causal video generation to perform coarse-to-fine multi-step reasoning before streaming output. On a six-task, level-stratified benchmark HDR improves success and progress markedly versus streaming autoregressive baselines, runs at 0.70 s/latent, and shows strong low-data robustness. Summary is based on the paper abstract.

Authors: Zezhong Qian, Xiaowei Chi, Chak-Wing Mak...
36 Twitter/X 2026-06-01 2 min read
Open

Andrew Ng (post published 2026-06-01) identifies the AI Forward Deployed Engineer…

Why it matters

Andrew Ng (post published 2026-06-01) identifies the AI Forward Deployed Engineer (FDE) as a rising Silicon Valley role: engineers embedded in client organizations to build and tune agentic workflows and customize off‑the‑shelf LLMs.

  • Ng argues there will be far more AI Engineer jobs than FDEs: companies may accept a few vendor‑embedded FDEs but will hire many more internal AI Engineers; he says his organizations hire far more AI Engineers than FDEs.
  • Ng notes FDEs were pioneered by Palantir ~20 years ago and require technical, communication, and business skills (client interviewing, strategy, explaining tech, and pushing back on unrealistic requests).
  • Ng warns vendor lock‑in is a common client concern—FDEs often deeply integrate specific vendors, reducing optionality—and sees surging demand for AI Engineers skilled in LLM prompting, agentic frameworks, evals, and AI coding agents (Claude Code, Codex, Antigravity CLI, OpenCode); he expects the AI Engineer role to fragment into specializations like LLMOps, Evals Engineers, AI Data Engineers, and Harness Engineers.

Andrew Ng argues that the AI Forward Deployed Engineer (FDE)—an engineer embedded inside client organizations to customize agentic workflows and adapt off‑the‑shelf LLMs—is resurging, but will remain a smaller slice of the market than the broader AI Engineer workforce. He traces the FDE model to Palantir about two decades ago and emphasizes FDEs’ mix of technical, communication, and business skills. Ng contends most companies will prefer building internal AI engineering capacity for vendor neutrality and long‑term optionality rather than relying heavily on vendor‑tied FDEs. He sees immediate, growing demand for AI Engineers proficient in LLM prompting, agent frameworks, evals, and modern coding agents (Claude Code, Codex, Antigravity CLI, OpenCode) and predicts the role will fragment over the next decade into specialized positions (LLMOps, Evals Engineers, AI Data Engineers, etc.), creating many new jobs rather than a job collapse.

By @AndrewYNg
37 ArXiv 2026-07-16 1 min read
Open

Stigmergic Graph Memory: An Environment-Aware Approach for Many-to-Many Multi-Agent Pickup and Delivery

Why it matters

Stigmergic Graph Memory (SGM) is a bounded, decaying memory layer that records recent execution signals on warehouse nodes and directed edges to rank feasible endpoints and route preferences; it does this without changing collision constraints or planner validity.

  • In experiments on five layouts × three load levels (15 map-load conditions) with 25 random seeds per condition, SGM beat two reconstructed many-to-many allocation baselines in all conditions, producing paired throughput gains of 20.5%–36.7%.

The paper tackles many-to-many Multi-Agent Pickup and Delivery (MAPD) where requests specify SKUs rather than fixed endpoints. It proposes Stigmergic Graph Memory (SGM), a decaying, environment-aware memory on nodes and edges that biases which feasible goals enter the planner (and how they are ranked) without altering planner validity. Across five layouts, three load levels, and 25 seeds per condition, SGM improved throughput by 20.5–36.7% versus two baselines. Summary based on the provided abstract and metadata; full text was not reviewed.

Authors: Aditya Dutta, Joon-Seok Kim
38 ArXiv 2026-07-16 1 min read
Open

HoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven Reasoning

Why it matters

Defines two quantitative metrics—Bias Intensity (BI) and Bias Harmfulness (BH)—and releases the LandmarkBias-3K benchmark (3,000 images) to measure the effect of landmark-induced bias on vision-language geo-localization models.

  • Introduces HoloGeo, an evidence-driven reasoning framework trained with BF-30k (30,000 images) annotated with structured multi-evidence, bias-free reasoning chains and multi-dimensional rewards; HoloGeo preserves performance on IM2GPS3K and YFCC4k while significantly outperforming open-source VLMs on LandmarkBias-3K.
  • Authored by Pengcheng Zhou, Xuanyu Liu, Yanchen Yin, et al.; posted to arXiv on 2026-07-16 (arXiv:2607.15255v1) with a PDF available at the provided link.

HoloGeo targets landmark bias in image geo-localization by proposing two metrics (BI, BH) and the LandmarkBias-3K benchmark to quantify bias effects. It trains an evidence-driven model using a new BF-30k dataset of structured, multi-evidence bias-free reasoning chains and multi-dimensional rewards to balance attention. According to the abstract, HoloGeo maintains IM2GPS3K/YFCC4k accuracy and outperforms open-source VLMs on LandmarkBias-3K. (Based on abstract.)

Authors: Pengcheng Zhou, Xuanyu Liu, Yanchen Yin...
39 ArXiv 2026-07-16 1 min read
Open

SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

Why it matters

SciDiagramEdit builds a benchmark that mines before/after figure pairs from arXiv version histories and operates on the figure's editable vector source so users can inspect and co-edit individual primitives.

  • The paper introduces agentic learning via skill evolution: an agentic proposer refines the agent's skill specification from execution traces over multiple epochs, which progressively increases edit accuracy on a held-out validation set.
  • Metadata: authored by Yasheng Sun, Zezi Zeng, Yifan Yang, Chong Luo, Wenyi Wang, Ziwei Liu, and Jürgen Schmidhuber; arXiv:2607.15272v1 (published 2026-07-16); 20 pages, PDF available.

SciDiagramEdit presents a benchmark and agentic skill-evolution framework for automated editing of scientific figures by learning from real arXiv paper revisions. The system works on vector figure sources and uses an agentic proposer that refines skill specifications from execution traces over epochs; this training on natural author edits yields progressively higher edit accuracy on held-out validation, demonstrating that paper revisions are an effective supervision signal.

Authors: Yasheng Sun, Zezi Zeng, Yifan Yang...
40 Twitter/X 2026-07-16 1 min read
Open

@alexocheema says he cannot share the original Steve Jobs Theatre slides, so he…

Why it matters

@alexocheema says he cannot share the original Steve Jobs Theatre slides, so he described one slide and used AI to generate a recreation.

  • The slide lists per-device local AI inference maxima: iPhone/iPad — up to 16GB unified memory and up to 14 billion active parameters; MacBook Air — up to 32GB and 35 billion; Mac mini — up to 64GB and 70 billion; MacBook Pro — up to 128GB and 120 billion; Mac Studio — up to 512GB and 480 billion; Mac Studio Cluster — up to 2TB unified memory and 1.6 trillion+ active parameters.

Alex Ocheema recreates a Steve Jobs Theatre slide he can't share, showing Apple devices' maximum local-AI inference capacities: iPhone/iPad up to 16GB and 14B active parameters; MacBook Air 32GB/35B; Mac mini 64GB/70B; MacBook Pro 128GB/120B; Mac Studio 512GB/480B; Mac Studio Cluster 2TB/1.6T+ parameters.

By @alexocheema
41 Twitter/X 2026-07-16 1 min read
Open

The idea for local.ai originated at Apple Park and, per @alexocheema on…

Why it matters

The idea for local.ai originated at Apple Park and, per @alexocheema on 2026-07-16 21:46:03+00:00, was built out at NVIDIA HQ using NVIDIA-provided hardware with 'no strings attached.'

  • local.ai is an independent benchmarking website for Local AI that benchmarked 1,000+ unique setups end-to-end with real agent harnesses, covering 'every hardware, every model, every quantization, every MTP setting,' and mapped results into Pareto frontiers with all data available free.
  • Access is being rolled out gradually; the author invites comments for early access and said the second batch of access codes would be sent 'in a few hours' from the 2026-07-16 post.

local.ai is an independent Local AI benchmarking site that, according to @alexocheema (2026-07-16), was conceived at Apple Park and built at NVIDIA HQ on NVIDIA-supplied hardware. The team ran 1,000+ end-to-end benchmarks with real agent harnesses across hardware, models, quantizations and MTP settings, visualized as Pareto frontiers, and is publishing the full dataset for free while rolling out staged early access.

By @alexocheema
42 Twitter/X 2026-06-30 4 min read
Open

Andrew Ng (post dated 2026-06-30) frames 'loop engineering' — a buzzphrase…

Why it matters

Andrew Ng (post dated 2026-06-30) frames 'loop engineering' — a buzzphrase amplified after viral mentions by Boris Cherny (Claude Code) and Peter Steinberger (OpenClaw) — as central to building 0-to-1 products.

  • Agentic coding loop: AI agents can write, test, and iterate until code meets a specification; Ng recounts a recent weekend where his coding agent worked autonomously for about one hour, cycling builds/tests every few minutes while he built a typing app for his daughter.
  • Developer feedback loop: humans review and steer the agent on timescales of tens of minutes to hours, translating vision into specs, updating them after implementations, and injecting a 'context advantage' (Ng's preferred term over 'taste'); persistent failures justify building evals.
  • External feedback loop: slow signals — friend feedback, alpha releases, A/B tests — take hours to weeks, inform developer vision, and, combined with faster coding agents, are moving engineers into partial product-manager roles.

Andrew Ng outlines three complementary 'loops' that now guide 0-to-1 product creation in the era of agentic coding, arguing that loop engineering — highlighted recently by Boris Cherny and Peter Steinberger — is shaping what and how we build. The agentic coding loop lets AI agents write, test, and iterate autonomously (Ng cites a real example where an agent ran for ~1 hour and produced builds every few minutes while he made a typing app for his daughter). The developer feedback loop operates on tens-of-minutes-to-hours cadences, where humans inject domain context (Ng calls this a 'context advantage') by shaping specs, adding evals, and steering product choices. The external feedback loop — friends, alpha testers, A/B tests — is slower (hours to weeks) but crucial for informing vision. Ng concludes that these loops together accelerate development and push engineers toward partial product-management responsibilities.

By @AndrewYNg
43 ArXiv 2026-07-16 1 min read
Open

Quantifying Training Membership Information in the Hyperspherical Embedding Geometry of Face Recognition Models

Why it matters

Study used a factorial design over IResNet backbone size, loss head, training duration, and number of training identities to train 180 face‑recognition models and quantify cluster-geometry membership signals.

  • Evaluated on nine benchmarks, the number of training identities produced the largest effect on member/non‑member separability; backbone and loss head contributed far less, and on a same‑domain held‑out reference the geometric membership signal decreased monotonically as more identities were added.
  • Cross‑domain non‑member sets (pose, age, quality, ethnicity) inflate the apparent membership signal, and fusing four cluster‑geometry statistics with a learned classifier reveals additional membership information beyond the best individual statistic.

The paper quantifies how hyperspherical embedding geometry in face‑recognition models encodes training membership. Using 180 models varying IResNet size, loss head, training duration, and number of identities, and evaluating on nine benchmarks, the authors find training‑identity count drives the strongest member/non‑member separability, same‑domain signals shrink as identities increase, cross‑domain sets inflate signals, and a learned fusion of four geometry statistics improves detection.

Authors: Ünsal Öztürk, Sébastien Marcel
44 Twitter/X 2026-07-16 1 min read
Open

Tori Shivanandan, Radiant President & COO, says Radiant is building microreactors…

Why it matters

Tori Shivanandan, Radiant President & COO, says Radiant is building microreactors and that 'small details' are critical to proving the product and de‑risking it for customers.

  • Shivanandan warns that when attempting novel engineering you 'can't assume the details will sort themselves out' and describes Radiant as 'detail obsessed' in addressing those minutiae.
  • She claims Radiant is 'closer than we’ve ever been' and 'further along than the headlines suggest' (post published 2026-07-16), invoking John Salvatier’s basement‑stairs essay as an analogy for unexpected technical details.

Tori Shivanandan, Radiant’s President & COO, says the company is building microreactors and that minute engineering details — often overlooked — are critical to proving and de‑risking their product. She argues you cannot assume details will resolve themselves, claims Radiant is 'closer than ever' (July 16, 2026), and uses a John Salvatier essay as an analogy.

By @a16z
45 Twitter/X 2026-05-14 1 min read
Open

Andrew Ng announced on 2026-05-14 a new course titled 'Transformers in Practice'…

Why it matters

Andrew Ng announced on 2026-05-14 a new course titled 'Transformers in Practice', built in partnership with AMD and taught by Sharon Zhou.

  • The course covers transformer internals: token-by-token generation, how attention and layers combine to predict tokens, and reasons LLMs hallucinate plus mitigations such as RAG and chain-of-thought.
  • It emphasizes hands-on diagnostics and deployment: diagnosing inference bottlenecks and GPU speedups (e.g., quantization) through interactive visualizations rather than video-only content.

The 'Transformers in Practice' course, announced by Andrew Ng on 2026-05-14 and taught by Sharon Zhou in partnership with AMD, offers hands-on exploration of transformer LLM internals—token-by-token generation, attention and layer interactions, causes of hallucination and mitigations like RAG and chain-of-thought, plus GPU inference diagnostics and speedups (e.g., quantization) via interactive visualizations.

By @AndrewYNg
46 Twitter/X 2026-07-17 1 min read
Open

Rebuilt a FedEx-facing delivery orchestration platform in 3.5 months with 2…

Why it matters

Rebuilt a FedEx-facing delivery orchestration platform in 3.5 months with 2 engineers, producing 122 merged pull requests in the first 90 days vs. the client's 7–8 month rebuild estimate.

  • Workflow: created a repository knowledge graph before writing code; ran a six-step ticket loop (define, spec, plan, implement, test, document); agents produced ~90% of code while senior engineers enforced a V.U.E. gate (verify, explain, debug without the agent).
  • Impact and cost: achieved ~50% time reduction due to upfront context, agent handling of mechanical coding, and strict spec-to-PR discipline; AI compute ≈ $200 per developer per month; CTO labeled the pod 'top performing team' and both engineers received discretionary bonuses twice.

Mark Ajzenstadt describes a two-engineer pod that rebuilt a FedEx-facing delivery orchestration platform in 3.5 months (122 merged PRs in 90 days) using a Velocity Framework: build a repo knowledge graph, run a six-step ticket loop, let agents generate ~90% of code, and gate merges with a V.U.E. senior-review requirement; AI compute cost ≈ $200/dev/month.

By @alex_prompter
47 ArXiv 2026-07-16 1 min read
Open

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning

Why it matters

AlphaWiSE is a post-hoc weight-space interpolation method that composes two frozen source checkpoints by fitting one scalar interpolation coefficient per aligned parameter tensor (the scalar is shared across all entries of that tensor).

  • The interpolation coefficients are fitted on a small exemplar memory to materialize a single interpolated checkpoint; the deployed model preserves the original architecture and parameter count and incurs no extra inference-time cost.
  • On audio-image-text retrieval tasks, AlphaWiSE shows consistent improvements over strong continual-learning baselines across multiple retrieval directions and evaluation metrics (ArXiv preprint by Sarthak Jain et al., posted 2026-07-16).

AlphaWiSE introduces a post-hoc weight-space interpolation technique for continual multimodal representation learning that composes two frozen checkpoints into one. For each aligned parameter tensor it fits a single scalar interpolation coefficient (shared across entries) on a small exemplar memory, producing an interpolated checkpoint with no additional inference cost. Experiments on audio-image-text retrieval report consistent gains over strong continual-learning baselines. Full text not available here (abstract-only).

Authors: Sarthak Jain, Qiran Hu, Zhen Zhu...
48 Twitter/X 2026-04-30 2 min read
Open

Karpathy presented three concrete 'new horizons' for LLMs at Sequoia Ascent 2026…

Why it matters

Karpathy presented three concrete 'new horizons' for LLMs at Sequoia Ascent 2026 (talk published 2026-04-30): (1) menugen — an app wholly driven by LLMs that takes an image and outputs an image with no classical code required; (2) installing '.md' skills instead of '.sh' scripts — writing installations in English for an LLM to interpret, target, and debug for a user's setup; (3) LLM knowledge bases that enable computation over unstructured data from arbitrary sources (text/articles), a capability he says was impossible with classical code.

  • Karpathy highlighted LLM 'jaggedness': the same model can coherently refactor a 100,000-line codebase yet also produce nonsensical directions (e.g., 'walk to the car wash'). He attributes this to domain verifiability plus economics — revenue/TAM shapes what frontier labs include in training distributions and RL packaging, so models are 'on the rails' when in-distribution and 'off-roading' otherwise.
  • He framed an emerging agent-native economy: products decomposed into sensors, actuators and logic across computing paradigms, a rising 'agentic engineering' skill set and new hiring practices, and speculative moves toward mostly neural computing augmented by classical CPU coprocessors.
  • Context and reaction: the remarks were part of a fireside chat ~a week before 2026-04-30; Karpathy (@karpathy) argued agentic engineering changes what can be built, and Stephanie Zhan (@stephzhan) contrasted last year's 'vibe coding' (raising the floor) with agentic engineering (raising the ceiling) and said Karpathy feels 'more behind as a programmer.'

Andrej Karpathy used a Sequoia Ascent 2026 fireside chat (talk published 2026-04-30) to argue that LLMs are creating qualitatively new capabilities, not just accelerating existing workflows. He gave three examples: 'menugen' (image→image apps fully handled by LLMs without classical code), replacing install .sh scripts with human-readable .md 'skills' the LLM executes and debugs, and LLM knowledge bases that perform computation over unstructured text from arbitrary sources. He also tackled LLM 'jaggedness' — why the same model can refactor a 100,000-line codebase yet hallucinate trivial actions — blaming a mix of domain verifiability and economic choices that shape RL training distributions. Finally, Karpathy sketched an 'agent-native economy' (sensors, actuators, logic), the rise of agentic engineering and hiring shifts, and the possibility of mostly neural computation aided by CPU coprocessors; Stephanie Zhan highlighted that this era raises the ceiling for what engineers can build.

By @karpathy
49 Twitter/X 2026-07-15 1 min read
Open

@swyx claims GPT-5.6 + Superapp now outperforms prior CUA efforts he tracked…

Why it matters

@swyx claims GPT-5.6 + Superapp now outperforms prior CUA efforts he tracked: World of Bits (Shi et al. 2017), Adept/@jluan (interview ~3 years ago), Anthropic's Computer Use launch (~2 years ago), Claude Cowork (~3 months ago), and a full CUA track at @aidotengineer (3 weeks ago).

  • He instructed his nontechnical team to use CUA for real knowledge‑work: signing up for payment and invoicing portals and handling speaker/sponsor/attendee/vendor/union data requests, asserting these tools are already fast and practical for everyday workflows.
  • He warns that anyone who agrees with a criticized take (he names Dwarkesh) is 'so not up to date' and that underestimating CUA capabilities is a 'dangerous category error' for AI decision‑making; he stresses this is a critique of one take, not of Dwarkesh personally.

swyx recounts following CUA from World of Bits (2017) through Adept/@jluan, Anthropic's Computer Use launch, Claude Cowork, and a recent ai:dot engineer track—and asserts GPT-5.6+Superapp now surpasses those systems. He has his nontechnical staff using CUA for real administrative workflows and warns that underestimating current capabilities (he critiques one take by Dwarkesh) is dangerous for AI decision-making.

By @swyx
50 ArXiv 2026-07-16 1 min read
Open

When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space

Why it matters

Hidden-state analyses across Qwen2.5-3B/7B/14B/32B, Phi-3.5, and SmolLM2 show content danger (CD) and physical danger (PD) form separable signals; PRISM (single-layer L2-regularized logistic probe over full hidden states) achieves 86.2–87.7% accuracy with 11.7–13.7% FPR on SafeAgentBench, versus 24.7–39.0% FPR for same-scale LLM judges.

  • On the new PSB-1K contrastive benchmark (1,000 physical-risk pairs without explicit harm keywords), PRISM reaches 99.6% accuracy and 0.7% FPR, while a Qwen2.5-3B judge wrongly rejects 67.8% of safe tasks; PRISM’s findings also replicate on SafeText and EARBench.

When Words Are Safe But Actions Kill evaluates whether physical danger (PD) and content danger (CD) are distinct signals in LLM hidden states. Using hidden-state direction analysis and random-split null tests across Qwen2.5-3B/7B/14B/32B, Phi‑3.5, and SmolLM2, the authors introduce PRISM — a single‑layer L2‑regularized logistic probe — and demonstrate strong PD detection (SafeAgentBench: 86.2–87.7% accuracy; PSB‑1K: 99.6%).

Authors: Weimeng Wang, Ziqiang Wang, Zihang Zhan...
51 Twitter/X 2026-07-16 1 min read
Open

Runta raised a $20M seed round led by a16z, announced on 2026-07-16.

Why it matters

Runta raised a $20M seed round led by a16z, announced on 2026-07-16.

  • Runta is rebuilding the execution layer for AI agents: a CPU-focused, extremely efficient runtime that can run locally or in the cloud and enforces security and policy controls to constrain agents while they run.
  • Founder Guanlan Dai (formerly led Cloudflare's edge proxy and Kong's core proxy) leads Runta, and Martin Casado is joining the Runta board.

Runta announced a $20M seed round led by a16z on July 16, 2026. The company, led by Guanlan Dai (ex‑Cloudflare/Kong), is building a CPU-focused execution layer for AI agents that prioritizes extreme efficiency, local or cloud deployment, and runtime security and policy controls to constrain agent behavior; Martin Casado joins the board.

By @a16z
52 Twitter/X 2026-07-17 1 min read
Open

@emollick (2026-07-17) flags Kimi K3 as having no model card and an expected…

Why it matters

@emollick (2026-07-17) flags Kimi K3 as having no model card and an expected weight release “in a couple weeks,” asking how pre-clearance would apply to open-weight models that are easy to jailbreak.

  • He asserts K3 is not yet at 'Mythos/Sol' capability but predicts someone will reach that level soon, increasing risk from open models.
  • He notes governments (US/UK/China) cannot recall downloaded weights but can force companies doing business with their citizens to avoid unvetted models, and argues this creates a need for international cooperation on model vetting.

Author @emollick raises pre-clearance questions for open-weight LLMs, citing Kimi K3's lack of a model card and an expected weight release “in a couple weeks.” He warns open models are easy to jailbreak, predicts Mythos/Sol-level models will appear soon, and says governments can’t recall downloads but can bar companies from using unvetted models, so international vetting cooperation is needed.

By @emollick
Quick skim

Scan these for facts, links, or weak signals worth tracking.

37 items · open
1 Twitter/X 2026-07-13 1 min read
Open

Richard S.

Why it matters

Richard S. Sutton announced on 2026-07-13 that he and Khurram Javed have broken away from John Carmack (@ID_AA_Carmack) and Keen Technologies to found a new startup, Oak Lab (@oaklab_ai).

  • Oak Lab affirms reinforcement learning and run-time experience as the basis for intelligence but claims current deep learning methods are "weak and inefficient" and require fundamentally new ideas and a thorough reworking (not mere tweaks) to support more ambitious AI goals; Sutton positions Oak Lab's path as distinct from Keen and Ineffable.

Richard S. Sutton announced on 2026-07-13 that he and Khurram Javed have left John Carmack's Keen Technologies to start Oak Lab (@oaklab_ai). Oak Lab endorses reinforcement learning and run-time experience as central to intelligence while arguing current deep learning is "weak and inefficient" and needs fundamental reworking, not incremental fixes, to reach ambitious AI objectives.

By @RichardSSutton
2 Twitter/X 2026-07-09 2 min read
Open

Adam Thierer warned on 2026-07-09 that the U.S. currently lacks clear statutory…

Why it matters

Adam Thierer warned on 2026-07-09 that the U.S. currently lacks clear statutory AI frameworks, leaving an opaque, sporadic de facto model-review/licensing regime in which 'voluntary' agreements and national-security pressure can force large model developers to comply or face market removals, long delays, or lost government contracts.

  • Andrew Ng (quoting Thierer on 2026-07-09) argues protecting open-source AI is essential to 'permissionless innovation,' warning that a 'restrict-until-permitted' presumption or direct national-security restrictions would effectively doom open-source models, recreating a Clipper Chip–era battle while open-source developers have no 'golden shares' to placate authorities.

Adam Thierer (quoted by Andrew Ng on 2026-07-09) warns the U.S. AI governance process is sliding toward an opaque, de facto model-licensing regime—driven by national-security officials and 'voluntary' agreements—that could impose pre-vetting or a 'restrict-until-permitted' rule. He likens the stakes to the 1990s Clipper Chip fight and urges a defense of open-source AI before it is effectively banned or crippled.

By @AndrewYNg
3 Twitter/X 2026-07-17 1 min read
Open

The ruptured pipe was a 100-year-old, 36-inch trunk line under Sunset Boulevard…

Why it matters

The ruptured pipe was a 100-year-old, 36-inch trunk line under Sunset Boulevard that LADWP had scheduled for replacement in 2031; LADWP budgeted $280 million to replace 6.4 miles of this exact trunk line, with design ongoing and construction planned 2031-2035—the pipe burst five years before crews would start.

  • Last fiscal year LADWP aimed to replace 1.1 miles of trunk pipe but replaced less than one; Los Angeles has roughly 550 miles of trunk line, so at the department’s new target of 3.5 miles/year each mile would be replaced every 157 years.
  • Riveted steel trunk pipe has an expected life of about 100 years; a 1921 trunk line ruptured in 2014 near UCLA, releasing roughly 20 million gallons and submerging Pauley Pavilion—today’s collapse is the same street and vintage, illustrating a gap between material lifespans and political replacement schedules.

The 100-year-old, 36-inch water trunk beneath Sunset Boulevard—already on LADWP’s list for replacement with a $280M, 6.4-mile project scheduled for 2031-2035—ruptured five years before planned work, creating a sinkhole. LADWP replaced <1 mile last year versus a 1.1-mile goal; at 3.5 miles/year the 550-mile system sees a 157-year cycle, outlasting riveted-steel’s ~100-year life.

By @aakashgupta
4 Twitter/X 2026-07-16 1 min read
Open

Aakash Gupta (tweeted 2026-07-16) asserts AI has collapsed the defenders' time…

Why it matters

Aakash Gupta (tweeted 2026-07-16) asserts AI has collapsed the defenders' time lag: attacker research that used to take days now takes minutes, so speed of scanning is no longer a sustainable edge—"whoever ships clean wins."

  • Prevention must move upstream: continuously rebuild container images so CVEs don't accumulate and attach SBOMs plus SLSA provenance to know exactly what is shipping.
  • Dan Lorenc (@lorenc_dan) warns "Open source is dying" and says his company Chainguard raised over $800M to build infrastructure that stops cyberattacks before they start.

AI-driven tooling has erased the traditional defender time advantage, Aakash Gupta warns (2026-07-16): exploit research now takes minutes rather than days, so detection and faster scanners won't suffice. He calls for upstream prevention—continuous image rebuilds, SBOMs and SLSA provenance—while Dan Lorenc says Chainguard raised $800M to protect open source.

By @aakashgupta
5 ArXiv 2026-07-16 1 min read
Open

Goal-Oriented Semantic Communication for Distributed ISAC-Enabled Vehicle Coordination

Why it matters

Introduces a goal-oriented semantic communication (GSC) framework for distributed ISAC-enabled vehicle coordination at unsignalized intersections; GSC transmits sensing and C&C signals only when semantically important for improving intersection throughput.

  • Uses an extended Kalman filter (EKF) to predict and fuse distributed RSU sensing, and a masked hybrid proximal policy optimization (MHPPO) that jointly selects sensing/C&C transmission decisions and C&C contents using a value-of-information (VoI) reward.
  • Adds an uncertainty-aware transmission design (UTD) — robust beamforming plus VoI-based time-division power allocation — and demonstrates in simulations (13 pages, 9 figures) 100% collision-free coordination with significantly reduced signaling overhead versus predictive ISAC baselines and ablations.

Liu and Deng propose a goal-oriented semantic communication framework for distributed ISAC vehicle coordination at unsignalized intersections, where multiple RSUs under a central BS collaboratively sense and send command-and-control (C&C). The system uses EKF for state fusion, a masked hybrid PPO (MHPPO) optimizing VoI-driven transmission and C&C content, and an uncertainty-aware transmission design (robust beamforming and VoI-based time-division power allocation). Simulations report 100% collision-free coordination and much lower signaling overhead than predictive ISAC baselines.

Authors: Wenjie Liu, Yansha Deng
6 ArXiv 2026-07-16 1 min read
Open

Catch, Throw, Repeat: Planning for Human-Robot Partner Juggling

Why it matters

Presents a real-time planning and control architecture that combines predictive ball tracking, adaptive online trajectory optimization using a multiple-shooting formulation, and a state-machine coordination logic to enable synchronized multi-ball human-robot partner juggling.

  • In an 8-participant user study (beginners to experts) the system achieved shared three-ball cascades; all participants surpassed previously reported best-case results within a 10-minute session. One participant extended the prior record for shared three-ball cascades fivefold to 20 consecutive robot catches; another achieved 100% success with 40 consecutive catches in a single-ball catch-and-return task.
  • Paper (Lippert, Ploeger, Chowdhury, Müller, Peters, Kshirsagar) posted on arXiv (2607.15129v1) and accepted to IROS 2026; project video: https://kai-ploeger.com/partner-juggling, PDF: https://arxiv.org/pdf/2607.15129v1

Human-robot partner juggling addresses dynamic object exchange under perception, timing, and contact uncertainty. The paper proposes a real-time system combining predictive ball tracking, adaptive online trajectory optimization via multiple-shooting, and state-machine coordination to synchronize multi-ball patterns with a human. In an 8-person study it produced reliable shared three-ball cascades and substantial record improvements, demonstrating a practical advance for physical human-robot interaction and shared autonomy.

Authors: Jonathan Rainer Lippert, Kai Ploeger, Abir Chowdhury...
7 ArXiv 2026-07-16 1 min read
Open

ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors

Why it matters

ARMOR++ (2026-07-16) achieves a statistically confirmed, substantial gain in blind-target Attack Success Rate (ASR) over the state-of-the-art agentic baseline on the AADD-2025 benchmark, and also outperforms non-agentic baselines and defended detectors.

  • The framework uses Qwen2.5-VL (vision-language model) to provide spatial semantic priors and Qwen3 (LLM) to orchestrate primitive selection, adaptive hyperparameter reparameterization, and entropy-regularized perturbation mixing.
  • ARMOR++ composes five complementary primitives—dense optimization, saliency-based methods, spatial transformations, frequency-domain perturbations, and block-structured modifications—to improve black-box, no-query transferability from convolutional surrogates to transformer-based deepfake detectors across low- and high-quality images.

ARMOR++ is an agentic multi-agent framework that improves black-box, no-query transfer attacks on deepfake detectors by combining Qwen2.5-VL spatial priors and a Qwen3 orchestrator to mix five complementary primitives (dense, saliency, spatial, frequency, block) with adaptive hyperparameter reparameterization and entropy-regularized mixing. Evaluated on AADD-2025, it significantly outperforms prior agentic and non-agentic baselines, revealing a persistent reliability gap in deployed detectors.

Authors: Christos Korgialas, Gabriel Lee Jun Rong, Dion Jia Xu Ho...
8 ArXiv 2026-07-16 1 min read
Open

Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation

Why it matters

Proposes a self-evolving, expert-in-the-loop annotation framework for Major Depressive Disorder (MDD) that pairs LLM-assisted labeling with expert verification and operates in three stages: candidate evidence selection, DSM-5-TR criterion-level analysis, and case-level synthesis; outputs labels plus clinical evidence, reasoning traces, and edit histories.

  • Introduces a dual-memory architecture (Example Memory and Reflection Memory) to internalize expert feedback and iteratively improve annotations without retraining; a pilot on expert-reviewed samples (arXiv:2607.15202v1, published 2026-07-16; accepted at IEEE COINS 2026) reportedly improved annotation consistency and explainability while reducing manual revision effort.

The paper presents a workflow to build explainable, DSM-5-TR-aligned depression annotation datasets by combining LLM-assisted candidate evidence selection, criterion-level DSM-5-TR analysis, and case-level synthesis with expert verification. A dual-memory (Example and Reflection) stores feedback to evolve annotations without model retraining. A pilot on expert-reviewed samples showed improved consistency, greater explainability, and reduced manual edits; full text available on arXiv.

Authors: Hoang-Loc Cao, Van Pham, Truong Thanh Hung Nguyen...
9 Twitter/X 2026-05-20 1 min read
Open

Andrew Ng posted on 2026-05-20 a new deeplearning.ai short course (partnered with…

Why it matters

Andrew Ng posted on 2026-05-20 a new deeplearning.ai short course (partnered with Google Cloud) taught by Katie Nguyen and Wafae Bakkali on building AI agents that generate images and videos, emphasizing agents that evaluate and iterate on their own outputs.

  • The course teaches three evaluation techniques combined in an agent: image-text similarity scoring, an LLM judge that scores against custom criteria like brand consistency, and structured rubrics of verifiable yes/no checks (e.g., "is the subject in the frame?", "does the camera motion match?").
  • Hands-on skills include image and video prompt engineering; building an image agent that turns brand guidelines into UI mockups; and building a video agent that plans multi-scene explainers and animates reference frames with synchronized audio (registration link: deeplearning.ai/courses/ai-a…).

Andrew Ng announced on 2026-05-20 a deeplearning.ai short course (with Google Cloud) taught by Katie Nguyen and Wafae Bakkali on building AI agents that generate images and videos by self-evaluating and iterating. The course covers three evaluation methods—image-text similarity, an LLM judge for brand criteria, and structured yes/no rubrics—and teaches prompt engineering, UI mockup image agents, and multi-scene video agents with synchronized audio.

By @AndrewYNg
10 Twitter/X 2026-07-15 1 min read
Open

@levie (published 2026-07-15) argues code is especially amenable to agents…

Why it matters

@levie (published 2026-07-15) argues code is especially amenable to agents because it can be quickly tested—either by manual checks or by running automated tests—giving fast feedback loops that most other work lacks.

  • Most non-code work only gets validated when it hits the real world (examples: a stock trade executes, a contract is negotiated, a sales pitch is delivered), so it lacks interim testability.
  • As a result, more agents will be layered into workflows and enterprises that build rigorous evals for knowledge work will capture the greatest AI gains; evaluation will become critical to agent adoption over time.

Levie contends that code’s rapid testability makes it unusually suitable for agent automation, whereas other domains lack immediate validation until real-world outcomes occur (e.g., trades, contracts, pitches). He predicts increased agent layering in workflows and stresses enterprises must develop stronger evals for knowledge work—those that do will benefit most from AI adoption.

By @levie
11 Twitter/X 2026-07-17 1 min read
Open

Detroit recorded an AQI of 724 on 2026-07-17; the 'Hazardous' category starts at…

Why it matters

Detroit recorded an AQI of 724 on 2026-07-17; the 'Hazardous' category starts at 301 and Delhi often registers around 400, and Berkeley Earth’s conversion (22 µg/m³ PM2.5 ≈ one cigarette/day) implies Detroit’s peak PM2.5 exposure equated to roughly a pack of cigarettes for a day spent outdoors.

  • The smoke source was 858 wildfires across Canada (113 officially 'out of control'); plumes traveled about 800 miles from Ontario to Michigan. Canada’s boreal forest covers 552 million hectares (larger than India), and many remote fires are monitored rather than actively fought because crews, aircraft, and access are insufficient—suppression for many relies on winter.
  • This is the third summer in four years that Great Lakes cities have topped global pollution rankings; 2023 burned 18.5 million hectares (Canada’s worst season). The boreal is drying rapidly and some 'zombie fires' survive underground over winter and reignite, shifting severe urban air-quality problems back to distant forest smoke despite decades of urban emission controls.

Detroit’s air reached an AQI of 724 on July 17, 2026, driven by smoke from 858 Canadian wildfires (113 out of control) whose PM2.5 traveled ~800 miles from Ontario. At peak, particulate exposure equated roughly to a pack of cigarettes per day; this marks the third summer in four years Great Lakes cities have led global pollution charts as a drying boreal produces persistent, winter-surviving 'zombie' fires.

By @aakashgupta
12 ArXiv 2026-07-16 1 min read
Open

Mutable Low-Rank Sketches for Retrain-Free Recommendation

Why it matters

Mutable sketches store each user in a KP-tree (a sparse segment tree with sum aggregation), fit one low-rank projection, and recompute embeddings on-the-fly; Theorem 1 shows each new observation monotonically tightens the prediction-error envelope (a guarantee FunkSVD and eALS lack).

  • On KuaiRec the mutable sketch achieves 0.810 RMSE while reading 1.8% of the data versus ALS at 0.822 RMSE using 100% of data; per-batch updates are 8× faster and a new user gets personalized recommendations in <1 ms after their first rating with no retraining.
  • KP-tree norm-proportional sampling yields 40–130% better item coverage on very sparse matrices (<1% density), while uniform sampling is adequate for dense matrices.

Mutable Low-Rank Sketches use a KP-tree to maintain sparse per-user preference sketches and a single low-rank projection to produce embeddings on arrival, enabling retrain-free updates. The paper proves monotonic tightening of the error envelope (Theorem 1) and reports strong KuaiRec results (0.810 RMSE at 1.8% read, 8× faster updates). Only the abstract was available.

Authors: Hector J. Garcia, Nick Clayton
13 ArXiv 2026-07-16 1 min read
Open

Decoding Market Emotion from Blockchain Activity: A Data-Driven Sentiment Classifier

Why it matters

XGBoost was the best-performing classifier in the study, achieving an average F1-score of about 0.84 for daily Twitter-based Bitcoin sentiment labels using cross-validation (paper published to arXiv 2026-07-16).

  • The model input merged on-chain transaction metrics, historical Bitcoin prices, and Twitter sentiment; SHAP was used to quantify on-chain feature contributions and the dataset was normalized for integrated analysis (accepted to ISCC 2026).

Decoding Market Emotion introduces a data-driven classifier that explains Bitcoin market sentiment by combining on-chain transactions, historical price data, and daily Twitter sentiment labels. The authors test multiple ML models and find XGBoost most reliable (avg. F1 ≈ 0.84 via cross-validation). They apply SHAP for feature-level interpretability, emphasizing explanation over price prediction and demonstrating meaningful signals for crypto market analysis.

Authors: Arthur G. Bubolz, Abreu Quevedo, Giancarlo Lucca...
14 Twitter/X 2026-07-17 1 min read
Open

Kimi-K3 scored 1679 points to reach #1 in Arena.ai's Frontend Code Arena…

Why it matters

Kimi-K3 scored 1679 points to reach #1 in Arena.ai's Frontend Code Arena (announcement posted July 17, 2026), overtaking Claude Fable 5.

  • Kimi-K3 jumped 17 places from Kimi-k2.6 (#18 → #1) and ranked #1 in 6 of 7 Frontend domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools; it placed #2 in Gaming behind Fable 5.
  • Author warns Arena ELO scores are limited and frontend chat UIs are easy to optimize with training/system prompts (saying people are "overindexing" on Arena scores); full Kimi-K3 model weights are promised for release by July 27, 2026.

Kimi K3 achieved a headline win on Arena.ai—1679 points and #1 in the Frontend Code Arena on July 17, 2026—jumping 17 places from Kimi-k2.6 and topping six of seven frontend domains. The author cautions that Arena ELOs are limited and chat front-ends can be tuned to be subjectively preferred; full weights arrive July 27, 2026.

By @emollick
15 Twitter/X 2026-07-17 1 min read
Open

On 2026-07-17 Guillermo Rauch announced Pete Hunt (@floydophone) has joined…

Why it matters

On 2026-07-17 Guillermo Rauch announced Pete Hunt (@floydophone) has joined Vercel to run Frameworks and lead Next.js; Hunt is described as an early React pioneer at Meta who powered Instagram Web adoption.

  • Nick Schrock (@schrockn), co‑inventor of GraphQL, joined Vercel to work on 'Agentic Developer Experience' aiming to enable the 'next billion agents' and build 'self‑improving software.'
  • Rauch framed the hires as a major moment ('HOLY FUCKING SHIT the react avengers have assembled'), called it a dream for a founder, and noted Vercel is hiring with the new hires' DMs open for applications or bug reports.

Guillermo Rauch announced on 2026-07-17 that Vercel has hired Pete Hunt and Nick Schrock. Hunt, a React pioneer at Meta who led Instagram Web adoption, will run Frameworks and lead Next.js. Schrock, co‑inventor of GraphQL, will work on Agentic Developer Experience to enable the 'next billion agents' and 'self‑improving software.' Rauch called the moment a dream and said they're hiring with open DMs.

By @swyx
16 Twitter/X 2026-07-16 1 min read
Open

An 11-year-old girl in rural Tennessee was reported missing; a nearby Flock…

Why it matters

An 11-year-old girl in rural Tennessee was reported missing; a nearby Flock street camera captured a car whose license plate matched a registered sex offender. Police used the plate and the vehicle's direction of travel on I‑75 to locate and rescue the bound but alive girl.

  • The speaker—identified as the founder of Flock, @glangley—said he started Flock nine years ago to solve cases like this; he recounted the story in a talk posted by @a16z on 2026-07-16 at @TEDTalks. Investigators later found at the suspect's home items 'to not only assault, but dispose of the body.'

A TEDTalk anecdote (posted by @a16z on 2026-07-16) describes a grandmother in rural Tennessee whose 11-year-old granddaughter vanished; a Flock street camera yielded a license plate tied to a registered sex offender. Officers tracked the car down I‑75, pursued and rescued the bound girl alive. The Flock founder, @glangley, said he launched the company nine years ago to prevent such crimes.

By @a16z
17 ArXiv 2026-07-16 1 min read
Open

ESAR: Event-Based Synthetic Aperture Reconstruction

Why it matters

ESAR formulates monocular event-camera reconstruction as a synthetic-aperture inverse problem that recovers a static ground-domain log–radiance field θ ∈ ℝ^{N_g}, replacing the latent pixel-time volume v ∈ ℝ^{N_pN_t} with the geometric relation v = Pθ and the linearized measurement model APθ = b + η (A = temporal differencing, b = signed binned event counts).

  • Under near-nadir motion successive projections are approximately shifted views so the composite operator AP is ill-conditioned (spatial averaging combined with temporal differencing); the authors (Antil, Blauvelt, Sayre; arXiv 2026-07-16) use regularized inversion and show on simulated data and real Falcon Neuro near-nadir event recordings that the θ-based method recovers coherent large-scale spatial structure while suppressing fine-scale texture relative to dynamic latent-image and learned event-reconstruction baselines.

ESAR (Event-Based Synthetic Aperture Reconstruction) frames monocular event-camera imaging as recovering a static log–radiance field θ via APθ = b + η, where P maps the scene into motion-dependent views and A is temporal differencing. Exploiting near-nadir motion (approximate shifts) and regularized inversion, experiments on simulated and Falcon Neuro data (Antil et al., 2026-07-16) recover coherent large-scale structure and suppress fine texture. Summary based on the abstract (full text not reviewed).

Authors: Harbir Antil, Daniel Blauvelt, David Sayre
18 ArXiv 2026-07-16 1 min read
Open

SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment

Why it matters

SUFLECA scales geometry-grounded feature learning with Normalized Object Coordinates (NOCs) supervision on 674,000 images spanning 12 real and synthetic datasets, producing compact geometry-aware features that generalize across domains.

  • A geometrically consistent matching algorithm establishes reliable one-to-one CAD-to-image correspondences for zero-shot 9D pose (rotation, translation, anisotropic scale) estimation, enabling sub-second alignment per object without iterative pose refinement.
  • On ScanNet25k SUFLECA achieves 33.4% category and 42.3% instance accuracy, outperforming the strongest zero-shot baseline by 10.3 and 12.2 percentage points respectively, and (reported) for the first time surpassing fully supervised methods; code: https://github.com/snt-arg/SUFLECA

SUFLECA addresses single-image CAD-to-image alignment by scaling up geometry-grounded feature learning and introducing a geometrically consistent matching algorithm. Trained with NOCs supervision on 674K images from 12 datasets, it learns compact, domain-general geometry-aware features and yields sub-second, zero-shot 9D pose alignment without iterative refinement. On ScanNet25k it reaches 33.4%/42.3% category/instance accuracy, beating prior zero-shot and reported fully supervised baselines. (Summary based on abstract only.)

Authors: Saad Ejaz, Miguel Fernandez-Cortizas, Javier Civera...
19 ArXiv 2026-07-16 1 min read
Open

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA

Why it matters

A retrospective GI endoscopy case study on MediaEval Medico 2025 compared nine documented multimodal VQA systems and found parameter-efficient adaptation of pretrained backbones produced the strongest challenge performance, but improvements in answer-level metrics did not consistently reflect faithful or complete clinical reasoning.

  • Models enforcing structured reasoning and explicit visual–textual grounding exhibited more reliable behavior across heterogeneous question types; the analysis is correlational rather than ablation-based and motivates evaluation beyond lexical overlap, standardized evidence-linked explanations, leakage-aware data governance, and lightweight robustness and calibration checks.

Healthcare multimodal AI for GI endoscopy is evaluated via a retrospective analysis of nine systems from MediaEval Medico 2025. Based on the abstract (full text not provided), parameter-efficient adaptation of pretrained backbones gave top challenge performance, yet did not guarantee faithful clinical reasoning. Methods with structured reasoning and explicit grounding were more reliable, prompting recommendations for stronger evidence-linked evaluation and governance.

Authors: Sushant Gautam, Vajira Thambawita, Michael A. Riegler...
20 ArXiv 2026-07-16 1 min read
Open

Motion-Conditioned Multi-View Fusion for Myocardial Infarction Localization from Echocardiography

Why it matters

MCF-Net achieves 72.4% F1 and 84.9% accuracy for segment-level myocardial infarction (MI) localization, outperforming motion-only, vision-only, and prior fusion baselines on echocardiography.

  • MCF-Net fuses EchoPrime foundation-model visual features from dual views with motion cues derived from extremely sparse supervision (a single annotated template frame for point tracking); motion-derived segment-aware soft masks and a motion-conditioned fusion module address view-dependent ambiguity (notably apical views).

MCF-Net is a motion-guided multi-view fusion framework for localizing myocardial infarction from echocardiography. It uses a pretrained EchoPrime foundation model across dual views and models cardiac motion with only a single annotated template frame to initialize point tracking. Motion-derived segment-aware soft masks and a motion-conditioned fusion module improve segment-level localization, yielding 72.4% F1 and 84.9% accuracy, surpassing motion-only, vision-only, and prior fusion methods.

Authors: Guang Yang, Wentian Xu, Siyu Wang...
21 ArXiv 2026-07-16 1 min read
Open

QuReC: All-in-One Image Restoration with Query-Specific Guidance and Local-Global Response Calibration

Why it matters

QuReC (Zhou et al., accepted at ACM MM 2026; arXiv:2607.15097v1, published 2026-07-16) introduces two modules: a Degradation-Guided Query Reconstruction Module (DQRM) that matches each spatial query to a degradation prototype space to produce query-specific degradation-aware representations, and a Local-Global Response Calibration Module (LGRCM) for dual-branch aggregation calibrated by learnable priors.

  • The authors add a weakly supervised prototype matching learning strategy to stabilize query-wise matching and improve degradation semantic consistency; experiments report that QuReC achieves superior performance on multiple all-in-one image restoration benchmarks and the code is released at https://github.com/zhoushen1/QuReC.

QuReC targets all-in-one image restoration under spatially heterogeneous and mixed degradations by combining per-query degradation-aware guidance with robust feature aggregation. DQRM reconstructs query-specific representations via matching to a learned degradation prototype space, stabilized with weakly supervised prototype matching, while LGRCM fuses local and global responses and calibrates them with learnable priors. The model outperforms prior unified restoration methods on multiple benchmarks; code and pretrained resources are publicly released.

Authors: Shen Zhou, Jinghui Zhang, Wenbo Huang...
22 ArXiv 2026-07-16 1 min read
Open

DriftWorld: Fast World Modeling through Drifting

Why it matters

DriftWorld learns an action-conditioned drift to generate multi-step future frames in a single forward pass, running at 30+ fps and producing rollouts on average 17× faster than diffusion-based world-model baselines.

  • On vision-based robotic benchmarks (Bridge-V2, RT-1, Language Table, Push-T, Robomimic) DriftWorld attains state-of-the-art decision-making performance while using far less inference time than diffusion models.
  • DriftWorld can be used as an offline simulator to rank real-world robot policies, with rollout-based scores correlating with ground truth up to 0.99.

DriftWorld introduces an action-conditioned drifting generative model that, instead of iterative denoising, learns a drift to produce multi-step visual rollouts in one forward pass. It runs at 30+ fps and is ≈17× faster than diffusion baselines, delivering state-of-the-art planning on Bridge-V2, RT-1, Language Table, Push-T, and Robomimic; offline ranking correlates up to 0.99.

Authors: Susie Lu, Haonan Chen, Weirui Ye...
23 Twitter/X 2026-07-16 2 min read
Open

30 engineering teams entered the company's internal hackathon; a PM (Jyothi…

Why it matters

30 engineering teams entered the company's internal hackathon; a PM (Jyothi Nookula) won by implementing an adversarial-agent approach inspired by a public Anthropic blog post and shipped it into production (post published 2026-07-16).

  • Architecture: one 'writer' agent generates outputs while a second, evaluator agent—configured with a precise, machine-checkable company-specific spec—attacks those outputs; flaws route back and the writer revises until the work clears the bar (described as 'GANs applied to agents' and demoed in Claude Code).
  • Gupta's claim: the builder agent and models/context window were commodity across teams; the scarce, decisive input was the written config/spec defining what 'good' means—PMs who own the evaluator/spec win.

Aakash Gupta recounts how a PM (Jyothi Nookula) beat 29 other teams in a 2026 hackathon by using adversarial agents (writer + evaluator) from an Anthropic-inspired design: the evaluator encodes a precise, machine-checkable company spec, attacks outputs, and forces iterative fixes until production quality is met. Gupta argues the spec/evaluator—not the model—is the real leverage.

By @aakashgupta
24 ArXiv 2026-07-16 1 min read
Open

Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence

Why it matters

Rubrics on Trial is a query-only framework that evolves a rubric set from an empty seed using only synthetic rubric-conditioned response pairs (no human-written rubrics, preference data, or model training); it validates each proposed rubric and filters out non-discriminative, over-specific, and style-only candidates.

  • In experiments across five preference benchmark suites, the method achieves the best average accuracy and leads on six of seven evaluation sets (ArXiv preprint 2607.15092v1, published 2026-07-16; authors: Haocheng Yang et al.).

Rubrics on Trial introduces a procedure that iteratively grows and validates query-specific rubrics from scratch by generating and checking synthetic rubric-conditioned response pairs, avoiding external annotations or model retraining. The framework screens out rubrics that fail to discriminate quality or that merely encode style, and empirically attains top average accuracy, winning six of seven evaluation sets across five preference benchmark suites (Yang et al., 2026).

Authors: Haocheng Yang, Licheng Pan, Xiaoxi Li...
25 Twitter/X 2026-07-16 1 min read
Open

Simon Dedic (@sjdedic) posted on 2026-07-16 that his profile-pic change is a…

Why it matters

Simon Dedic (@sjdedic) posted on 2026-07-16 that his profile-pic change is a strategic signal marking the end of an era: after starting his crypto career in 2017 he intends to move away from pseudo-anonymous/comic pfps and embrace doxxing/accountability at Moonrock Capital.

  • He credited degen culture with building today's ecosystem but argued the next decade will be won by fundamentals, real track records and people 'willing to sign their name under their conviction,' ending with 'Costume's off. Fundamentals up.'

Simon Dedic of Moonrock Capital announced on 2026-07-16 that a new profile picture is a strategic shift away from pseudo-anonymity—after a crypto career begun in 2017 he wants doxxed leadership, verifiable track records and accountable teams. He praised degen culture’s role but insisted future success requires fundamentals, measurable value creation and public conviction.

By @0xdasha
26 Twitter/X 2026-07-14 1 min read
Open

On 2026-07-14, @fchollet praised Harvinder and Suman’s Airtap for turning SMS…

Why it matters

On 2026-07-14, @fchollet praised Harvinder and Suman’s Airtap for turning SMS into a headless, agentic execution layer that operates mobile apps (explicitly naming DoorDash and TikTok) in the background while providing plain-text updates and only prompting the user for authentication.

  • Airtap is reachable at airtap.ai and via SMS at +1 (650) 213-7322; the product “watches your apps, escalates until the goal lands,” chains workflows across phones (cloud or on-device), and uses the tagline: “Set an Airtap. Get your life back.”

@fchollet tweeted on 2026-07-14 that Harvinder and Suman’s Airtap turns SMS into a headless agentic execution layer that runs apps like DoorDash and TikTok in the background, chains multi-step workflows across phones (cloud or on-device), returns plain-text progress, and requires user intervention only for authentication; site airtap.ai and SMS contact +1 (650) 213-7322.

By @fchollet
27 ArXiv 2026-07-16 1 min read
Open

Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies

Why it matters

The authors compare Grokipedia (released late 2025) and Wikipedia using 1,394 article pairs about government members, scored along nine expert-coded ideological dimensions and judged by four LLMs: Grok, Claude, Mistral, and DeepSeek (paper posted 2026-07-16).

  • All four LLM judges rate Grokipedia as less neutral than Wikipedia; both encyclopedias portray politicians favourably overall, but Grokipedia favours economically right-wing politicians and penalises socially liberal ones, whereas Wikipedia is rated as favourably biased toward socially liberal politicians.

The paper evaluates political neutrality in Grokipedia versus Wikipedia by analysing 1,394 paired articles on government members across nine expert-coded ideology dimensions, using four LLM judges (Grok, Claude, Mistral, DeepSeek). All judges rate Grokipedia as less neutral; both sites portray politicians positively but with different slants—Grokipedia toward economically right-wing actors and against socially liberal ones, while Wikipedia leans favorable to socially liberal figures.

Authors: Filippos Vlahos, Guillaume Bied, Tijl De Bie
28 Twitter/X 2026-07-17 1 min read
Open

Argentina removed 13 zeros from its currency between 1970 and 1992; one modern…

Why it matters

Argentina removed 13 zeros from its currency between 1970 and 1992; one modern peso equals 10 trillion of the original pesos.

  • A MAX linear 'peso goes to zero' chart is misleading: the flatline is a resolution artifact — on a log scale the steep inflation phase ended about two years ago.
  • Key current metrics: annual inflation peaked near 300% in early 2024 and is now ~33.5%; the peso trades around 1,487 per USD with weekly volatility ~0.4%; currency bands were scrapped in April 2025 and the peso has floated under a $20B IMF program, though Argentines still price major assets in dollars.

Argentina's peso has been redenominated repeatedly (13 zeros removed between 1970–1992), so a long linear MAX chart exaggerates past collapses and hides recent dynamics. The author argues the steep inflation phase ended ~2024; inflation fell from ~300% peak to ~33.5%, the peso floats near 1,487/USD with ~0.4% weekly volatility under a $20B IMF program, but it remains a poor store of value and many prices persist in dollars.

By @aakashgupta
29 Garry's List 2026-07-16 13 min read
Open

The Merchant and the Lawyer

Why it matters

Sinclair Louie (1950s San Francisco) bought a Sea Cliff home with legal help from Jewish lawyer Ben Lehr, a symbolic case of Chinese–Jewish solidarity that prefigured multiethnic Civil Rights coalition work.

  • Federal data and research: a 2022 Federal Reserve breakout found median Asian American household net worth at $536,000 vs. $285,000 for white households; Brookings/Pew cited median Black wealth ~$44,900 and Hispanic ~$61,600, highlighting a fifty-year wealth divergence within the coalition.
  • San Francisco anti-Asian violence surged: 9 reported victims in 2020 to 60 in 2021 (a 567% increase); six Asian Americans were killed in San Francisco between 2020–2023; high-profile victims named include Yik Oi Huang and Vicha Ratanapakdee.
  • Grassroots political response: Asian voters led recalls and removals—Feb 2022 (three SF school board members), June 2022 recall of DA Chesa Boudin (Asian voters backed recall at 67%), Nov 2024 recalls in the East Bay removing Sheng Thao and Pamela Price; Sept 2025 SF Supervisor Joel Engardio was also recalled.

The article argues that Chinese and Jewish Americans built, funded, and supplied plaintiffs, lawyers, and activism for the Civil Rights coalition across the twentieth century—examples include Yick Wo v. Hopkins, Lau v. Nichols, the San Francisco State strike, Japanese internment redress, and organizations such as Chinese for Affirmative Action, the Asian Law Caucus, and AALDEF. Using concrete data (a 2022 Federal Reserve breakout showing median Asian American net worth of $536,000 vs. $285,000 for whites, and Brookings/Pew wealth figures for Black and Hispanic households), the author traces a fifty-year divergence in outcomes among coalition members.

He documents two recent episodes of abandonment: the 2020–2023 wave of anti-Asian violence (San Francisco reporting 9 victims in 2020 and 60 in 2021, six local fatalities from 2020–2023, and named victims like Yik Oi Huang and Vicha Ratanapakdee) and a 2024–2025 surge in antisemitism (ADL’s 2024 audit: 9,354 incidents; 2025 lethal attacks including shootings at the Capital Jewish Museum and a Boulder firebombing). When institutions on the left framed or minimized these harms—favoring de‑carceral explanations or selective narratives—the affected communities organized politically (recalls and electoral change) and built parallel protections. The piece closes by invoking the Sinclair–Lehr story as a model for renewed mutual solidarity and a practical call to support Jewish neighbors now.

By Forrest Liu
30 ArXiv 2026-07-16 1 min read
Open

MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos

Why it matters

MAGiSt3R is a multi-agent 3D reconstruction framework that processes monocular RGB videos at almost 10 FPS using a feed-forward 3R-family model to regress local point maps and a dedicated merging module (MAGMA) to fuse maps intra-agent and inter-agent.

  • The system applies pose-graph optimization to mitigate cumulative camera drift and — according to evaluations on synthetic and real-world datasets — obtains superior reconstruction and camera-tracking accuracy versus state-of-the-art (Gong et al., arXiv:2607.15211v1, 2026-07-16).

MAGiSt3R is a multi-agent, feed-forward 3D reconstruction system for monocular RGB videos that achieves nearly 10 FPS. It regresses local point maps with a 3R-family network, fuses them with the MAGMA merger at intra- and inter-agent levels, and reduces drift via pose-graph optimization. Authors report superior reconstruction and tracking on synthetic and real datasets; only the abstract was available.

Authors: Ziren Gong, Xiaohan Li, Fabio Tosi...
31 ArXiv 2026-07-16 1 min read
Open

Ray-based phase error correction for miniaturized DOE projector-based FPP under single-directional hyperbolic projection

Why it matters

Proposes a ray-based phase-error correction framework for miniaturized DOE projector-based Fringe Projection Profilometry (FPP) that models phase artifacts along projection rays from the projector pinhole, avoiding image-domain or neighboring-pixel processing.

  • Introduces a projector-pinhole estimation method using a single-directional hyperbolic fringe pattern to recover projector geometry without stereo calibration, and a data-efficient refinement built from a single calibration pose.
  • Authors Seung-Jae Son, Yatong An, and Jae-Sang Hyun (preprint 2026-07-16) report experiments showing significant improvements in reconstruction accuracy under nonlinear projection conditions for miniaturized DOE projector FPP systems.

Ray-based phase error correction for miniaturized DOE projector-based Fringe Projection Profilometry (FPP) targets severe phase artifacts caused by nonlinear projection and limited pattern control. The approach models artifacts along projector rays, estimates the projector pinhole from a single-directional hyperbolic fringe pattern to recover geometry without stereo calibration, and refines with a single calibration pose; experiments report significant accuracy gains.

Authors: Seung-Jae Son, Yatong An, Jae-Sang Hyun
32 ArXiv 2026-07-16 1 min read
Open

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Why it matters

Authors Paul Kassianik, Blaine Nelson, and Yaron Singer (arXiv 2026-07-16) propose a cost-aware evaluation that compares security agents at fixed cost levels and decomposes performance into inference spend versus tool (telemetry/enrichment) spend; results and an interactive site are at https://evals.frontier.security.

  • On offensive Cybench CTF challenges, performance improves with additional test-time compute; the paper reports that scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive.
  • On defensive Splunk BOTS v1 SOC investigations, success depends more on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning/inference budget, so economic efficiency and operational fit matter more than peak success rate.

Cost-aware evaluation of LLM-based security agents measures performance at fixed monetary/compute budgets and decomposes costs into inference and tool spend. The authors evaluate offensive Cybench CTFs and defensive Splunk BOTS v1 SOC investigations, finding offensive tasks scale with test-time compute—scaled open-weight models approach proprietary frontiers cost-competitively—whereas defensive success hinges on disciplined tool use, telemetry navigation, and selective enrichment rather than raw reasoning budget.

Authors: Paul Kassianik, Blaine Nelson, Yaron Singer
33 Twitter/X 2026-07-16 1 min read
Open

Kimi K3 (an open Chinese model) beat Opus 4.8 on automated retro-game generation…

Why it matters

Kimi K3 (an open Chinese model) beat Opus 4.8 on automated retro-game generation for Road Fighter, Battle City, and Q*bert, producing the best Q*bert where the player jumps on cubes, paints them, and flees a purple snake; Kimi reportedly got gameplay, physics, mechanics, and AI behavior all working together.

  • Token usage and cost per run: Kimi K3 used 18.4K tokens for $0.28, GPT-5.6 used 18.1K tokens for $0.28, and Opus 4.8 used 21.3K tokens for $0.54 — Kimi K3 cost almost 2× less than Opus 4.8.
  • GPT-5.6 struggled on cars and produced a broken Battle City (tank died, base got hit); the author frames Kimi K3 as an open Chinese model competing head-to-head with US ‘frontier’ labs backed by billions as of mid‑2026.

Kimi K3, an open Chinese model, beat Opus 4.8 in automated creation of three retro arcade games (Road Fighter, Battle City, Qbert), delivering the strongest Qbert and integrated gameplay/AI. Tests used identical prompts; Kimi used 18.4K tokens ($0.28) versus Opus 4.8’s 21.3K tokens ($0.54). The author presents this as mid‑2026 parity with well‑funded US labs.

By @adxtyahq
34 Twitter/X 2026-07-15 1 min read
Open

Ernesto Lopez (@ErnestoSOFTWARE) says his bootstrapped, two‑person app makes…

Why it matters

Ernesto Lopez (@ErnestoSOFTWARE) says his bootstrapped, two‑person app makes $50,000/week across iOS and Android — roughly $200,000/month — and ranks #7 against well‑funded competitors (post dated 2026-07-15).

  • Lopez lists a specific stack he calls non‑negotiable: Rork (Fable 5 for updates/betas), FunnelFox (web→app funnels; claims ~30% savings on Apple fees when running Meta ads), Higgsgield MCP (AI UGC and AI ad testing), Amplitude, Singular, Superwall, Claude Code, and Sideshift.
  • @alexcooldev counters that tool choice doesn’t matter to users — what matters is solving the user’s problem — and reiterates that distribution remains the most important skill.

Ernesto Lopez claims his two‑person, bootstrapped app earns $50,000/week (~$200,000/month) and credits a precise tool stack (Rork/Fable 5, FunnelFox, Higgsgield MCP, Amplitude, Singular, Superwall, Claude Code, Sideshift) for growth. @alexcooldev adds that users don’t care about tooling; solving the problem and strong distribution are what drive success.

By @alexcooldev
35 Twitter/X 2026-07-15 1 min read
Open

On 2026-07-15 Aaron Levie stated "Code is OP," arguing code's testability (you…

Why it matters

On 2026-07-15 Aaron Levie stated "Code is OP," arguing code's testability (you can run tests or manually verify behavior) makes it uniquely amenable to agent-driven automation.

  • Levie notes many non-code knowledge workflows only reveal outcomes once deployed — e.g., "a stock trade is executed, a contract is negotiated, a sales pitch is delivered" — so they lack immediate evalability.
  • He asserts enterprises that build robust evaluations for knowledge-work workflows and layer agents into processes will gain the most from AI; evaluation capability will become critical for agent adoption.

Aaron Levie (2026-07-15) argues that "code is OP" because code's fast testability makes it ideal for agent automation; most other knowledge work (e.g., a stock trade executed, a contract negotiated, a sales pitch delivered) lacks immediate evals. He warns enterprises must build better workflow evaluations and layer agents into processes to capture AI gains.

By @fchollet
36 ArXiv 2026-07-16 1 min read
Open

DAPGNet: Dynamic Adaptive Physics-Guided Graph Diffusion Network for Hyperspectral Image Classification

Why it matters

DAPGNet achieves state-of-the-art OA, AA, and Kappa on Indian Pines, WHU-Hi-LongKou, Houston2013, and Houston2018; it improves average accuracy (AA) by 3.64–7.31 percentage points over the strongest competing method.

  • Architectural innovations include node-wise multiscale physical-prior encoding of contiguous spectral responses; a two-stage prior-aware sparse graph constructor combining spectral-spatial affinity, physical-prior consistency, and spatial distance; learned edge weights converted to additive attention biases; a physical gate for node-/feature-wise interpolation between graph-aggregated and projected physical-prior features; cross-scale fusion; and training with main classification, auxiliary supervision, and second-order spectral smoothness regularization.

DAPGNet is a dynamic adaptive physics-guided graph diffusion network for hyperspectral image classification that injects contiguous-band physical priors into relation-level graph learning. It encodes multiscale spectral priors per node, constructs a prior-aware sparse topology, converts learned edges into attention biases, and uses a physics gate plus cross-scale fusion. On four benchmarks (Indian Pines, WHU-Hi-LongKou, Houston2013, Houston2018) it yields top OA/AA/Kappa, with AA gains of 3.64–7.31 points. The paper (ArXiv 2026-07-16) reports ablation and sensitivity studies validating each component.

Authors: Pengkun Wang, Weijia Cao, Ning Wang...
37 ArXiv 2026-07-16 1 min read
Open

Concept-Guided Spatial Regularization for World Models in Atari Pong

Why it matters

Reproduced five visual world-model agents (DreamerV3, DIAMOND, TWISTER, Simulus, STORM) on Atari Pong, froze each learned world model, and evaluated them with closed-loop rollouts driven by separately trained policies.

  • All five frozen models showed clear simulation failures (ball disappearance, incorrect ball motion, invalid ball–paddle interactions). Pixel-space zero-shot MBRL policies trained entirely inside frozen models underperformed their original agents—DreamerV3 mean return dropped from -5.5 to -20.9 (near Pong minimum -21).
  • Introduced Concept-Guided Spatial Regularization (CGSReg), an auxiliary pixel-reconstruction loss applied to segmented concept regions; CGSReg improved closed-loop rollouts and zero-shot MBRL for DreamerV3, DIAMOND, and TWISTER, but effects varied across models and metrics.

The paper diagnoses failures of five visual world models in Atari Pong by freezing reproduced models and testing them with closed-loop rollouts and pixel-space zero-shot MBRL. It finds pervasive visual/dynamical errors and large policy-performance drops (e.g., DreamerV3 from -5.5 to -20.9). The authors propose Concept-Guided Spatial Regularization (CGSReg), which boosts rollouts and zero-shot policy quality for several agents but is not a complete fix.

Authors: Yukuan Lu, Zaishuo Xia, Weyl Lu...