Brief
DexTwist introduces a functional twist-retargeting method for MR-based teleoperation that targets contact-rich rotational manipulation where kinematic imitation fails. The system detects tripod pinches, estimates intended screw axis and twist, then performs a real-time joint-space residual optimization minimizing a virtual-object objective (turn angle, axis consistency, fingertip closure, tripod stability). Simulations and real tests demonstrate improved turning-angle tracking and reduced screw-axis drift versus a vector-based baseline.
Authors: Dongmyoung Lee, Chengxi Li, Dongheui Lee
Brief
CUActSpot targets the long-tail of complex GUI interactions that undermine computer-use agents by providing a multimodal benchmark (GUI, text, table, canvas, natural image) and a renderer-based data-synthesis pipeline: automatic scene generation, screenshot/element-coordinate recording, and LLM-produced instructions/action traces. Training on this corpus yields Phi-Ground-Any-4B, which outperforms open models under 32B parameters. Only the abstract was available for this summary.
Authors: Miaosen Zhang, Xiaohan Zhao, Zhihong Tan...
Brief
A popular dining-focused credit card received an anniversary refresh on May 11, 2026 that adds a 5X hotel-earning rate, nearly $100 in limited-time travel and dining credits, and rental-car elite status (enrollable now). AwardWallet notes the upgrade is coupled with a welcome bonus up to 100,000 points, though the email withheld the issuer/card name for advertising reasons.
By AwardWallet
Brief
Task-Adaptive Embedding Refinement via Test-time LLM Guidance presents a test-time method that refines query embeddings via a generative LLM's feedback on a small document subset to tailor embeddings to ad-hoc zero-shot search and classification tasks. Experiments with state-of-the-art embedding models across diverse benchmarks report consistent gains (up to +25% relative), improving ranking and class separation; code released on GitHub. Abstract only; full text not provided here.
Authors: Ariel Gera, Shir Ashury-Tahan, Gal Bloch...
Brief
The paper introduces a proximal-gradient Monte Carlo sampler for composite log-concave densities π ∝ e^{-f-g}, combining gradient steps on smooth f with a restricted Gaussian-oracle (RGO) proximal sampler for g. Under α-strong convexity of f+g and β-smoothness of f it achieves ε total-variation error in ~O(κ·sqrt(d)·log^4(1/ε)) iterations (κ=β/α), and the authors extend guarantees to Poincaré/LSI targets and to Lipschitz non-smooth f.
Authors: Linghai Liu, Sinho Chewi
Brief
Ramp's merch fulfillment service (announced in an email published 2026-05-11) combines printing with inventory storage, pick-and-pack operations and worldwide shipping, plus online store setup, onboarding kits and event/campaign distribution. The offering is modular — clients can select full fulfillment or specific services — and was credited with selling out The Bugle Podcast's limited-edition Christmas jumpers after Ramp printed and shipped orders globally.
By Ramp for Merch
Brief
MoLA (Mixture of Latent Actions) targets the gap between video-based imagination and actionable control: instead of feeding predicted frames to a policy or decoding videos directly into controls, it infers a mixture of latent actions via pretrained, modality-aware inverse-dynamics models (semantic, depth, flow) to produce a physically grounded action interface. Evaluated on LIBERO, CALVIN, LIBERO-Plus and real robots, the abstract reports consistent improvements in success rates, temporal consistency, and generalization; summary based on the abstract (full paper not reviewed).
Authors: Yajie Li, Bozhou Zhang, Chun Gu...
Brief
EgoEV-HandPose tackles egocentric 3D bimanual hand-pose estimation and gesture recognition from stereo event cameras by introducing KeypointBEV, which lifts stereo features into a bird's-eye-view and iteratively reprojection-refines depth and kinematic estimates. Trained and evaluated on the new EgoEVHands dataset (5,419 sequences, 38 gestures), it reports MPJPE 30.54 mm and 86.87% Top-1 accuracy, outperforming RGB-stereo and prior event-based methods, notably under low-light and occlusion.
Authors: Luming Wang, Hao Shi, Jiajun Zhai...
Brief
Laplacian-based neural operators are analyzed for the generalized Gierer–Meinhardt reaction–diffusion system: the paper obtains explicit approximation-error bounds depending on network depth, width, and spectral rank by using the Laplacian eigenfunction expansion of the PDE Green’s function. The authors show parameter complexity scales at most polynomially with accuracy and present numerical experiments consistent with theory. Summary based on the abstract; full text not reviewed.
Authors: Takashi Furuya, Ryo Ozawa, Jenn-Nan Wang
Brief
The paper tackles low-latency, whole-body humanoid teleoperation by mapping Virdyn IMU suit data to a Unitree G1 using a custom motion-processing, kinematic-retargeting, and control pipeline that avoids offline buffering and learning-based modules. Validated in MuJoCo then transferred unchanged to the physical robot, the system reportedly achieves stable, synchronized reproduction of a wide motion repertoire; summary based on the abstract and metadata.
Authors: Hamza Ahmed Durrani, Suleman Khan
Brief
TMRL and Context-Smoothed Pre-training (CSP) inject forward-diffusion noise into policy inputs during pretraining to create a continuum from precise imitation to broad action coverage, then train agents to modulate the diffusion timestep during RL fine-tuning to control exploration. The method works with states, 3D point clouds, and visual policies and enables sub-hour real-world manipulation fine-tuning; full paper and code on arXiv and project site.
Authors: Matthew M. Hong, Jesse Zhang, Anusha Nagabandi...
Brief
The paper develops a model-based bootstrap for transition kernels in finite controlled Markov chains with possibly nonstationary or history-dependent policies, proving distributional consistency in both long-chain and episodic offline RL regimes. Using a novel bootstrap LLN and a martingale CLT, the authors extend results to OPE and OPR via Hadamard-differentiable Bellman operators, producing asymptotically valid CIs; RiverSwim experiments show strong empirical calibration.
Authors: Ziwei Su, Imon Banerjee, Diego Klabjan
Brief
The paper introduces the first online L2D algorithm for multiclass classification with bandit feedback and a dynamically varying expert pool, addressing streaming data and shifting expert availability. It achieves regret O((n+ne)T^{2/3}) generally and O((n+ne)√T) under low noise, relying on new H-consistency bounds and first-order online convex optimization; experiments validate practicality. Summary based on the abstract (full text not available).
Authors: Dang Hoang Duy, Yannis Montreuil, Maxime Meyer...
Brief
Siebenborn et al. formalize bilateral morphological symmetry in bimanual mobile manipulation and propose a C2-equivariant flow-matching policy that enforces reflection symmetry through loss regularization or an equivariant velocity network. On planar and 6-DoF tasks the method boosts sample efficiency and enables zero-shot generalization to mirrored states, with real-world TIAGo++ validation. Summary based on the abstract; full text not reviewed.
Authors: Max Siebenborn, Daniel Ordoñez Apraez, Sophie Lueth...
Brief
Optimal policy learning under combined budget and minimum-coverage constraints is treated as a knapsack-type allocation problem; Cerulli (May 2026) proves the optimal rule is an affine threshold in budget and coverage shadow prices, shows an LP relaxation has an O(1) integrality gap, and evaluates two algorithms (GLC, RC), with Monte Carlo confirming near-optimal finite-sample performance and predictable failure modes.
Authors: Giovanni Cerulli
Brief
TriBand-BEV introduces a fast LiDAR-only 3D pedestrian detector that encodes the full point cloud into a lightweight 2D BEV tensor with three height bands, reformulating 3D detection as 2D detection and reconstructing boxes post-hoc. Using area attention, a hierarchical bidirectional neck (P1–P4), and a distribution-focal rotated-IoU head, it reaches 58.7/52.6/47.2 BEV AP on KITTI at 49 FPS; code is public.
Authors: Mohammad Khoshkdahan, Alexey Vinel
Brief
TextSeal presents a practical, localized watermark for LLM outputs that combines Gumbel-max sampling, dual-key generation, entropy-weighted scoring, and multi-region localization to preserve diversity while enabling strong provenance detection. The method adds no inference cost, supports serving optimizations like speculative decoding, strictly dominates prior baselines (e.g., SynthID-text), is robust to dilution, transfers through distillation, and a 6,000 A/B multilingual study (5 languages) reported no perceptible quality change. Full paper text was not available in the provided content.
Authors: Tom Sander, Hongyan Chang, Tomáš Souček...
Brief
AmbiSuR revisits Gaussian Splatting to improve photometric‑ambiguity‑robust 3D surface reconstruction. The authors uncover two primitive‑wise ambiguities and an intrinsic self‑indication ability in the representation, then introduce photometric disambiguation and an ambiguity‑indication module to constrain and correct geometry. Experiments reportedly yield superior reconstructions across challenging scenes; paper on arXiv (2026-05-12) and accepted at ICML 2026.
Authors: Jiahe Li, Jiawei Zhang, Xiao Bai...
Brief
Multi-Variable Conformal Prediction (MCP) tackles the limitation of conventional conformal methods that use a scalar score and single threshold by allowing vector-valued scores and multiple calibration variables. Using scenario theory, MCP unifies prediction-set design and calibration into one optimization problem (no data split) and provides finite-sample coverage. Two variants, RemMCP and RelMCP, trade off convexity assumptions and conservatism; experiments on ellipsoidal and multi-modal sets show target coverage, smaller/comparable set sizes, and lower calibration variance. Full text on arXiv.
Authors: Laura Lützow, Simone Garatti, Marco C. Campi...
Brief
FuTCR (Future-Targeted Contrastive and Repulsive) tackles Continual Panoptic Segmentation by preventing the collapse of diverse unlabeled objects into a single background representation. The method groups predicted masks with background labels but non-background logits to find future-like regions, uses pixel-to-region contrast to form coherent prototypes, and repels background features from known-class prototypes. According to the abstract, FuTCR yields up to 28% relative gains on new-class panoptic quality while maintaining or improving base-class performance (up to 4%), evaluated across six CPS settings and varied dataset sizes.
Authors: Nicholas Ikechukwu, Keanu Nichols, Deepti Ghadiyaram...
Brief
MEME (Multi-entity & Evolving Memory Evaluation) targets LLM-agent failures when storing, updating, and reasoning about many entities across sessions. The benchmark defines six tasks (including Cascade, Absence, Deletion) and tests six memory systems across three paradigms on 100 controlled episodes. Results show catastrophic collapse on dependency reasoning (Cascade 3%, Absence 1%), and only an expensive file-based agent + Claude Opus 4.7 partially closes the gap, highlighting a practical-performance tradeoff.
Authors: Seokwon Jung, Alexander Rubinstein, Arnas Uselis...
Brief
EgoForce tackles depth–scale ambiguity and device-specific generalization in monocular, head-mounted hand capture by fusing a differentiable forearm model, an arm–hand transformer that predicts geometry from a single egocentric view, and a ray-space closed-form solver to recover absolute camera-space 3D pose. The method works across fisheye, perspective, and wide-FOV optics and yields up to 28% MPJPE reduction on HOT3D, with code and data released.
Authors: Christen Millerdurai, Shaoxiang Wang, Yaxu Xie...
Brief
The Algorithmic Caricature (Gunjan et al., arXiv 2026-05-12) evaluates whether LLM-generated political posts replicate real online populations by comparing a paired corpus of 1,789,406 posts across nine crisis events. It finds synthetic text is fluent but population-level unrealistic—more negative, less sentiment-dispersed, structurally regular, and lexically abstract—with gaps varying by event and summarized by a proposed 'Caricature Gap'. Full text not available; summary based on abstract.
Authors: Gunjan, Sidahmed Benabderrahmane, Talal Rahwan
Brief
OmniNFT (Zhang et al., arXiv 2026-05-12) targets RL fine-tuning for joint audio–video generation by diagnosing three failure modes—advantages inconsistency, gradient imbalance, and uniform credit assignment—and proposing modality-wise advantage routing, layer-wise gradient surgery, and region-wise loss reweighting in an online diffusion-RL pipeline. Tests on JavisBench and VBench with the LTX‑2 backbone report improved per-modality quality, cross-modal alignment, and synchronization. Summary based on the abstract.
Authors: Guohui Zhang, XiaoXiao Ma, Jie Huang...
Brief
The paper presents a real-world dataset aimed at enabling AI-native mobility in 6G by replacing common simulation-based data with measurements from a commercial network. To address high interruption times and measurement overhead during UE mobility, the authors collected multi-speed traces across pedestrian, bike, car, bus, and train scenarios, emphasizing handover events. A key contribution is inclusion of timing advance (TA) at RACH trigger, MAC CE, and PDCCH grant events. The authors provide dataset generation details and exploratory analyses and propose use cases such as TA prediction and AI/ML-driven beam and handover management (arXiv:2605.12453v1).
Authors: Mannam Veera Narayana, Rohit Singh, Deepa M. R...
Brief
Perception Deep Research frames open-world visual perception where target identities must be resolved from external web facts before localization. The authors introduce WebEye — a benchmark with 120 images, 473 annotated objects, 645 QA pairs and 1,927 task samples — and propose Pixel-Searcher, an agentic search-to-pixel workflow that attains top open-source results across grounding, segmentation, and VQA.
Authors: Bokang Yang, Xinyi Sun, Kaituo Feng...
Brief
Shen and Kuusela (2026) introduce a score-augmented loss for neural likelihood surrogates in simulation-based inference, augmenting binary cross-entropy with exact parameter-space score ∇_θ log p(x | θ) and adaptive weighting based on loss gradients. Evaluated on network dynamics and spatial processes, the approach boosts surrogate quality and can match the effect of 10× more training data with under a 10% training-time increase.
Authors: Alexander Shen, Mikael Kuusela
Brief
GuidedVLA proposes guiding Vision-Language-Action models by supervising individual attention heads with manually defined auxiliary signals, rather than relying on end-to-end implicit learning. The paper implements three specialized heads (object grounding, spatial geometry, temporal skill logic) and reports higher success rates on simulated and real-robot tasks versus strong VLA baselines. Full text was not available in the provided abstract.
Authors: Xiaosong Jia, Bowen Yang, Zuhao Ge...