ArXiv

AI-based single-shot structured-light depth reconstruction for real-time laparoscopic surgical guidance

Authors
Wayne Wonseok Rodgers, Xiangyi Le, Seonghoon Jang...
Categories
eess.IV, cs.RO, physics.optics
arXiv
https://arxiv.org/abs/2608.05109v1
PDF
https://arxiv.org/pdf/2608.05109v1

Brief

A synchronization-free, single-shot structured-light laparoscopic depth system projects a passive LED binary mask into one channel of a dual-channel scope and trains a VQ-VAE with a latent-space U-Net depth head on 722 Zivid-referenced phantom pairs. It achieved MAE 3.70 mm, AbsRel 0.0326, delta=1.1 0.962, and ran at 26.0 Hz on an NVIDIA A100, enabling video-rate endoscopic depth estimation without explicit segmentation.

Why it matters

Platform: a synchronization-free, single-shot structured-light laparoscopic system using a passive LED-illuminated binary mask coupled to one channel of a dual-channel laparoscope; training used 722 Zivid 3D-referenced phantom image pairs reprojected into the SSLE frame.

Key details

  • Quantitative results: the VQ-VAE encoder with a latent-space U-Net depth head achieved MAE = 3.70 mm, AbsRel = 0.0326, delta=1.1 accuracy = 0.962, and delta=(1.1)^2 accuracy = 0.970; it yielded lower MAE than a dual U-Net (MaskNet+DepthNet) baseline and outperformed off-the-shelf monocular depth models on MAE, AbsRel, and threshold accuracy.
  • Runtime and practical notes: the pipeline ran at 26.0 Hz over 301 consecutive frames on an NVIDIA A100 GPU; authors highlight the importance of dataset size and SSLE–Zivid calibration accuracy and report no separate segmentation/mask-prediction stage.
Source evidence

Abstract

Significance. Accurate intraoperative depth perception is important for autonomous and semi-autonomous robotic laparoscopic surgery. Conventional fringe projection profilometry can achieve millimeter-scale accuracy but often requires multi-shot acquisition, digital-micromirror-device projection, and projector-camera synchronization, complicating integration into compact laparoscopic systems. Aim. To develop a synchronization-free, single-shot depth-sensing platform using a passive LED-illuminated binary mask and a VQ-VAE prior with a custom U-Net depth head. Approach. A compact projection module was coupled to one channel of a dual-channel laparoscope, while the second channel imaged the fringe-illuminated target. A Zivid 3D camera acquired reference depth for 722 paired phantom images. Zivid depth maps were reprojected into the SSLE image frame for supervised training and evaluation. The VQ-VAE encoded each input into a discrete latent representation, and a latent-space U-Net predicted depth without a separate mask-prediction branch. Results. Using a fixed train/validation/test split, the proposed model achieved an MAE of 3.70 mm, AbsRel of 0.0326, delta=1.1 accuracy of 0.962, and delta=1.1^2 accuracy of 0.970. It achieved lower MAE than the dual U-Net MaskNet + DepthNet baseline and outperformed off-the-shelf monocular depth models in MAE, AbsRel, and threshold accuracy. The pipeline operated at 26.0 Hz over 301 consecutive frames on an NVIDIA A100 GPU. Conclusions. The LED-illuminated binary-pattern platform with latent-space depth reconstruction enables synchronization-free, video-rate endoscopic depth estimation. Results demonstrate Zivid-referenced phantom reconstruction without an explicit segmentation stage, while emphasizing the importance of dataset size and SSLE-Zivid calibration accuracy.

Comment: 17 pages, 7 figures