ArXiv

AI_LectureNote's post‑ASR workflow restores Latin‑script medical terms and raised…

Authors
Kyeongeon Lee, Donghoon Chang, Seungryeol Baek...
Categories
cs.CL
arXiv
https://arxiv.org/abs/2607.17237v1
PDF
https://arxiv.org/pdf/2607.17237v1

Brief

AI_LectureNote presents a retrospective, readability‑oriented post‑ASR workflow that rewrites Korean‑English medical lecture transcripts to restore Latin‑script medical terms instead of Korean phonetic transliterations. Evaluated on four author‑recorded lectures across five conditions, it substantially improved English‑script rendering (whisper‑1: 0.39→0.71; gpt‑4o‑transcribe 3‑min chunks: 0.26→0.65) but produced measurable semantic drift and polarity errors, motivating separate evaluation axes; this is a single‑annotator pilot.

Why it matters

AI_LectureNote's post‑ASR workflow restores Latin‑script medical terms and raised macro English‑script rendering from 0.39 to 0.71 on the whisper‑1 path and from 0.26 to 0.65 on 3‑minute chunked gpt‑4o‑transcribe output (evaluation on four author‑recorded lectures across five conditions).

Key details

  • Script restoration did not guarantee meaning preservation: the two post‑processed conditions showed semantic drift in 34 and 36 of 282 reference sentences and polarity failures in 11 and 13 of 101 polarity‑cue rows, respectively.
  • Failure patterns varied by type and front end: polarity‑failure overlap across front ends had Jaccard 0.60 (9 shared of 15 unioned failures) versus general semantic‑drift overlap Jaccard 0.23 (13 shared of 57 unioned drifts); study is a single‑annotator pilot and recommends separate evaluation of surface accuracy, term‑script rendering, chunk‑level script consistency, and medical‑meaning preservation.
Source evidence

Abstract

AI_LectureNote is a historical, readability-oriented post-ASR workflow for Korean-English medical lectures. It rewrites speech-to-text output into study transcripts while restoring Latin-script medical terms rather than Korean phonetic transliterations. We retrospectively evaluate the workflow on four author-recorded lectures across five conditions. In this pilot, post-processing raised the macro English-script rendering rate from 0.39 to 0.71 on the whisper-1 path and from 0.26 to 0.65 when applied to 3-minute chunked gpt-4o-transcribe output. However, English-script rendering did not imply semantic faithfulness: the two post-processed conditions showed semantic drift in 34 and 36 of 282 reference sentences and polarity failures in 11 and 13 of 101 polarity-cue rows. A descriptive cross-input comparison suggested different candidate failure patterns: polarity-failure sets overlapped more strongly across front-ends (Jaccard 0.60; 9 shared of 15 unioned failures) than general semantic-drift sets (Jaccard 0.23; 13 shared of 57 unioned drifts). This single-annotator pilot documents concrete failure modes rather than population rates and supports evaluating surface accuracy, term-script rendering, chunk-level script consistency, and medical-meaning preservation separately.

Comment: 12 pages, 4 figures