Abstract
Surgical video interpretation is a promising medical artificial intelligence application. However, no existing video annotation method preserves the spatiotemporal complexity of surgeon reasoning. Here we show that verbal reasoning and visual attention can be converted into structured, machine-actionable records of intraoperative behaviours. Our method decomposes transcribed verbal commentary into video-anchored semantic feedback chunks, which are classified via a large language model, with spatial grounding to surgical scenes via eyegaze or cursor tracking. We demonstrate method validity and scalability on structured and unstructured annotation tasks. For quality feedback on full-length colorectal procedures, the method reached near-human fidelity for chunking (cosine similarity: 0.95±0.01) and semantic classification across observations (Cohen’s κ: 0.71±0.07) and evaluative triggers (Cohen’s κ: 0.67±0.14), with excellent usability ratings. For structured critical view of safety assessment in laparoscopic cholecystectomy, implicit annotation yielded excellent agreement with explicit reviewer ratings (Cohen’s κ: 0.83, 0.49 and 0.81 across three criteria). We anticipate this method will advance surgical data science by enabling scalable construction of meaningfully annotated surgical video datasets.
Competing Interest Statement
F.R.K. declares an ongoing advisory role for Scopia AI, Canada, and has received research funding from Novartis. D.S. is a consultant for Johnson and Johnson and Applied medical and receives research support by Beckton Dickinson, Intuitive surgical, and Cook medical. The other authors declare no competing interests.
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
The details of the IRB/oversight body that provided approval or exemption for the research described are given below:
All data collection was approved by the Purdue University Institutional Review Board (IRB #2025-00000221). Before participation, informed consent was obtained. No identifying information was stored alongside the review data, audio recordings were deleted after transcription had been verified, and all data were held on Purdue University's secure file transfer protocol.
I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.
Yes
Data Availability
Source data and code are available from the corresponding author upon reasonable request. The large language models used for chunking and classification are publicly available commercial models named in the Methods.





