Employing Vision-Language Model for Structured Evaluation of Oral Surgery Videos: a Four-Level Multimodal Framework with Expert Validation
Abstract
Oral Surgery videos are rich but unstructured sources of clinical information. This report presents an end-to-end system that automatically turns oral surgery videos into structured, clinically meaningful evaluations. The system applies a four-level multimodal AI framework — visual perception and localization, sequential workflow modeling, multimodal clinical reasoning, and OSATS-compliant skill scoring — and pairs every AI prediction with a human-expert verification that includes a grading platform, discrepancy detection with third-party arbitration, and a medical safety checker. We describe the problem and social impact, objectives, a comparison with existing solutions, methodology, preliminary results, and a roadmap for future improvement.
1. Problem Statement & Social Impact
Every year, millions of surgical procedures are performed, and in teaching hospitals a growing share of them are recorded on video. Oral surgery is no exception: operations such as cyst enucleation, impacted‑tooth removal, and implant placement are routinely captured for training, documentation, and medico‑legal purposes. Yet almost all of this visual data is unstructured and unused. A typical video is watched, if at all, only by the surgeon who performed it or by a handful of trainees. There is no standard, scalable way to turn what happens in the operating room into objective, comparable, and reviewable assessments of what was done and how well it was done.
This gap matters for three reasons. First, manual review is slow and subjective. Assessing a single operation against a clinical standard can require an expert to re‑watch the video, annotate its phases, and judge instrument handling — work that takes longer than the operation itself and varies from one reviewer to another. Second, surgical training depends on feedback. Junior surgeons need to know not only whether they completed an operation but whether they omitted a clinical step, chose the right instrument, or handled tissue gently. Today such feedback is scarce and inconsistent. Third, beyond training and quality control, building reliable machine understanding of operative video is also a critical foundational prerequisite for next‑generation autonomous surgical robots. Before intelligent agents can plan, predict, or execute surgical maneuvers independently, they must be able to perceive surgical scenes, parse procedural workflows, and interpret technical performance from raw video inputs. Robust multimodal evaluation pipelines help benchmark this machine comprehension capability and lay groundwork for future robotic surgical development.
Oral surgery presents unique technical challenges that further complicate automated video analysis. Unlike many open‑field surgical disciplines, intra‑oral procedures take place within a confined, deep oral cavity with limited depth‑of‑field for camera capture. The field of view is frequently obscured by intermittent blood, irrigation mist, saliva, and rapidly moving handpieces and instruments. Tissue boundaries are often subtle, and viewpoints shift constantly as the surgeon adjusts angles for intra‑oral access. Publicly available high‑quality annotated operative video datasets for oral surgery remain scarce, which adds additional barriers for building robust computer‑vision models. These domain‑specific obstacles make automated interpretation of oral surgical footage considerably harder than general‑purpose surgical‑video tasks.
Meanwhile, medical artificial intelligence has advanced rapidly. Text‑based systems can now converse with patients and reason about clinical cases at an expert level, and recent work has extended this to real‑time video consultations [1]. However, analyzing an actual operative video — recognizing anatomy, reconstructing the surgical workflow, forecasting the next step, and scoring technical skill — remains largely unsolved, particularly for surgical specialties such as oral surgery that are underrepresented in mainstream AI research. We therefore address the following question: can a multimodal AI system automatically analyze oral surgery videos and produce a structured, clinically meaningful evaluation that experts can trust?
If successful, the benefits are substantial. Hospitals and dental schools could evaluate trainees objectively and at scale; expert surgeons could review many more cases per unit time; and junior surgeons anywhere in the world, including those without access to senior mentors, could receive standardized, actionable feedback. This work supports a broader vision in which AI augments — rather than replaces — human expertise in surgical education and quality improvement.

Figure 1. System overview. An oral surgery video is analyzed by a four-level multimodal AI framework (perception, workflow modeling, clinical reasoning, and OSATS skill scoring) into a standardized evaluation report, while an expert final stage validation — independent grading, discrepancy detection, third-party arbitration, and a safety checker — establishes trustworthy ground truth.
2. Objectives
This project aims to develop and pilot-validate an end-to-end multimodal AI system for automated structured evaluation of oral surgery videos, to enable scalable, standardized surgical training and quality assurance.
Specifically, this work seeks to:
- Design a four-level hierarchical evaluation framework (perception, workflow modeling, clinical reasoning, OSATS scoring) operationalized as a tiered prompt system for vision-language models.
- Build a curated, difficulty-graded, privacy-compliant oral surgery video dataset with standardized nomenclature and full provenance traceability.
- Establish an expert-reviewed validation pipeline including a web grading platform, dual independent review, automatic discrepancy detection, and third-expert arbitration.
- Define a layered quantitative evaluation protocol with domain-specific metrics for each framework level.
- Pilot-test the full end-to-end system, verify technical feasibility, and characterize initial performance and key improvement priorities.
3. Comparison with Existing Work
Several families of solutions already exist, and each differs from ours in task, performance, cost, and adaptability.
Text- and image-based medical AI. General-purpose large language models and image-based systems have shown expert-level performance in clinical conversation and in interpreting static images. The most relevant recent milestone is AMIE (Video) [1], which demonstrated expert-level AI performance in real-time clinical video consultations with simulated patients: clinical evaluators rated it on par with or better than primary care physicians on history-taking, diagnosis, and physical examination (an overall rubric score of 83% versus 68% for physicians), and its top-1 diagnostic accuracy reached 91% versus 77% for physicians. Yet AMIE analyzes consultations — conversations with patients — not surgical operations. Its authors explicitly note remaining limitations in fine anatomical precision, subtle affective nuance, and high-frequency movements, and it does not reconstruct operative workflow or score technical skill. Our work targets a different, complementary problem: the operation itself.
Automated surgical phase recognition. In other specialties (most often laparoscopy), deep-learning models have been trained to segment videos into surgical phases. These systems can be accurate on the phases they were trained on, but they typically recognize only a fixed, predefined set of events, require large annotated datasets, and are rarely combined with clinical reasoning or skill assessment. They also cannot explain why a phase was missed or predict what should happen next.
Human OSATS assessment. The Objective Structured Assessment of Technical Skills (OSATS) rubric is the clinical gold standard for evaluating surgical skill [2], but it is normally completed by a human observer in real time or from a recording. It is reliable when performed carefully, yet it is expensive, slow, and inherently subjective.
Finally, the curated oral-surgery-specific dataset with expert-validated labels and a difficulty gradient is itself a differentiating asset: most public oral surgery video datasets are small, single-procedure, and lack clinically meaningful annotation.
4. Methodology
4.1 Curated Oral Surgery Video Dataset
This project constructs a standardized, difficulty-graded corpus of oral surgery videos that serves both as training and validation material for the multimodal model and as a standalone clinical teaching resource. The dataset is built through purposive, criteria-based sampling rather than random selection, and the overall pipeline comprises three stages: source retrieval, tiered screening, and standardized annotation.
4.1.1 Video Source Identification and Collection
Three complementary strategies are used to cover publisher-hosted, platform-hosted, and institution-licensed operative content, ensuring both diversity and clinical authority.
- Publication-driven strategy. A two-stage search of PubMed and PubMed Central (PMC) is conducted across 142 dentistry journals indexed in the Clarivate Master Journal List (2016–2026). Stage 1 retrieves 27,235 unique procedure articles through 3,834 journal-by-journal queries, after automated exclusion of reviews, editorials, letters, commentaries, and news items. Stage 2 maps the records to PMC and scans full-text XML for embedded or linked video files. Articles with downloadable operative video supplements are retained, forming the publication-derived subset reported in the PRISMA flow diagram.
- Platform-driven strategy. Clinician-oriented video-sharing platforms and publisher video archives are systematically surveyed. MedTube is enumerated category by category across seven dental specialty sections; YouTube is searched using a controlled vocabulary of procedure terms developed from specialty-society terminology (AAE, EFP/AAP, AAOMS, ITI, and related sources); and the freely accessible archive of the Quintessence Dental Video Journal (video articles published 2008–2014) is enumerated in full. Source-level exclusion filters — removing lectures, news stories, interviews, consultations, and non-English material — are applied during enumeration.
- Institutional-resource strategy. Video content licensed through the University of Hong Kong library is reviewed, comprising the Alexander Street Dentistry collection (enumerated after excluding lecture-type and non-procedural document types via the collection's native filters) and ClinicalKey dental textbooks containing embedded clinical videos.
Every retained video is logged with its source platform, original URL, publisher, publication year, and title. Collection is performed programmatically using automated, resumable, checkpointed scripts, followed by file-integrity verification and cross-source duplicate detection. The final index comprises 1,114 videos from eight sources: MedTube (289), YouTube-hosted dental content (301), publication-derived video supplements (95), the Quintessence Dental Video Journal (122), the Alexander Street Dentistry collection (142), ClinicalKey textbook videos (61), and two supplementary instructional collections (104). Every video in the corpus can be traced to its origin through a structured provenance index (watch URL, publisher, year, licensing status).
4.1.2 Inclusion and Exclusion Criteria
Screening operates at two levels — source level and instance (video) level.
- Source-level eligibility. A source is retained only if it provides operative (rather than conceptual or lecture-based) content, English-language narration or text, and material that is either freely accessible on a public platform or licensed to the authors through institutional library subscription. Lecture recordings, news items, interviews, panel discussions, consultations, and non-procedural content are excluded at this stage, prior to instance-level screening.
- Instance-level inclusion criteria (all must be satisfied): the video depicts a genuine dental operative or surgical procedure performed on a human patient; the footage has stable focus; the footage has adequate lighting; the key procedural steps are visible; and the video shows a complete phase of a procedure or a complete procedure.
- Instance-level exclusion criteria (any one triggers exclusion): non-operative content, including lectures, conference presentations, teaching-only content without live operative footage, software-screen workflow demonstrations (e.g., CAD/CAM design walkthroughs), imaging-only or diagnostic-only content, and oral examination-only content without a subsequent surgical procedure; procedures not performed on real patients, including demonstrations on phantom heads, typodonts, or dental models, animation or 3D simulation, and cadaveric or animal experiments; visibility of the patient's whole face; identifiable patient information (name, medical record number, institution-specific patient identifier, date of surgery, subtitles, consent forms, charts, or information-bearing watermarks); duplicate or substantially overlapping videos relative to videos already retained; and operative-field obscuration exceeding 50% of the frame.
4.1.3 Annotation and Difficulty Gradient Design
A dental operation/procedure keyword library is built and used to drive an AI-assisted YouTube search that returns more than 5,000 candidate videos. After an initial pass over the first 1,000+ candidates and the application of the explicit inclusion/exclusion criteria above, roughly 500 videos are retained; additional downloads are bringing the dataset toward a target of approximately 2,000 videos.
Every video is annotated with its operation/procedure type using the framework's standardized nomenclature. On top of this, a difficulty-gradient classification of dental operations is established: each operation standard is assigned a difficulty level reflecting its technical requirements and clinical risk, so each video inherits the level of its operation type. This dual-purpose resource serves AI training — letting the model reason about a procedure's demands and risks — and education, giving students a risk-aware map from lower to higher difficulty.
4.2 Vision-Language Model (Google Gemini) Pipeline
At the core of the system is a native multimodal vision-language model (VLM) in the Gemini architecture, which supports unified sequence modeling of video, audio, and text and converts raw surgical video into structured clinical outputs through an end-to-end token stream.
4.2.1 Input Sampling and Tokenization
The input is a raw oral surgery video with an audio track. Preprocessing tokenizes the multimodal signals and aligns them in time:
- Visual channel: video frames are sampled at 1 FPS, with each frame generating 258 visual tokens (66 tokens in low-resolution mode).
- Audio channel: each second of audio generates 32 audio tokens.
- Sequence unification: temporal positional encodings are added, and visual and audio tokens are integrated into a single unified token sequence that fully preserves frame order and temporal information.
4.2.2 Model Architecture and Generation
The model core is a native multimodal Mixture-of-Experts (MoE) Transformer with a long-context window of approximately 1M tokens (corresponding to roughly one hour of video). Cross-modal attention aligns pixel, audio, and text semantics, and the model adaptively captures surgical actions and scene changes. Driven by the hierarchically designed prompt instructions described in Section 4.3, the model generates structured outputs autoregressively, covering object recognition, timestamped phase segmentation, clinical reasoning, and skill scoring.

Figure 2. Video frame extraction at 1 FPS → 258 visual tokens per frame (66 tokens for low‑resolution mode) → 32 audio tokens per second → unified token sequence formed with temporal positional encoding → native multimodal MoE Transformer (1M‑token context ≈ 1‑hour video) with cross‑attention for aligning pixels, audio and text → autoregressive generation of structured language outputs (recognition, timestamp‑aware phase segmentation, reasoning, skill scoring).
4.3 Four-Level Multimodal Clinical Assessment Framework
A four-level clinical evaluation framework is designed to decompose surgical video analysis into increasingly higher-order tasks. The framework is operationalized as a tiered prompt system for the VLM, with a strict structured-output template that pairs every AI prediction with a dedicated human-expert evaluation block.
Level 1 — Visual Perception and Localization
This layer answers: what is happening, where, and with what? The model identifies the type of operation being performed; lists the anatomical and pathological structures it can see using standard medical names (e.g., "maxillary second premolar" or "periapical cyst"); lists the instruments used in the order they appear; and specifies spatial orientation (left or right, and surfaces such as buccal, lingual, or palatal).
Level 2 — Sequential Workflow Modeling
This layer answers: what was done, in what order, and what was skipped? The model divides the operation into a timeline of phases, each with a start and end timestamp. It then compares this timeline with standard surgical protocols to list any clinical steps that were missed or omitted up to the current point in the video.
Level 3 — Multimodal Clinical Reasoning
This layer answers: is the surgery complete, and if not, what should happen next? If the operation is not finished at the end of the video, the model determines the current surgical state and predicts the exact, medically correct next action, together with its rationale. A built-in safety checker verifies that the recommendation is not dangerous or hallucinated and that medical terminology is used correctly — for example, flagging a predicted next step that contradicts standard surgical protocol.
Level 4 — Cognitive Alignment through OSATS-Based Skill Scoring
This layer achieves cognitive alignment with expert surgical judgment using the six dimensions of the OSATS rubric: respect for tissue, time and motion, instrument handling, knowledge of instruments, flow of operation and forward planning, and knowledge of the specific procedure. The model rates the surgeon from 1 (novice) to 5 (expert) on each dimension. Crucially, every score must be justified by a concrete, timestamped visual event — for example, "score 5 for tissue handling because the mucosa was retracted gently without tearing" — which sharply reduces hallucinated ratings.
4.4 Expert Grading Web Platform
A web platform is developed at https://grading-platform-for-oral-surgical.vercel.app/ to enable independent expert verification of AI outputs and to construct a reliable ground-truth generation pipeline. Its core functions include:
- Dual independent review. Two oral surgeons independently grade the AI outputs for the same video on the platform.
- Automatic disagreement detection. The system automatically detects where the two experts disagree and highlights those items.
- Third-party arbitration. When the two experts cannot agree, the item is routed to a third expert for arbitration, creating a reference standard that is not tied to any single reviewer's judgment.
- Embedded annotation protocol and ground-truth templates. The platform standardizes medical nomenclature and the scoring anchors for each OSATS dimension, making the review process reproducible.
- OSATS-dedicated interface. The review interface provides separate scoring fields for each of the six OSATS dimensions and explicitly records whether an AI scoring justification constitutes hallucination.
A key design choice is the deliberate separation of "fast and automatic" from "authoritative": the AI provides a fast, consistent first pass that can scale to hundreds of videos, while the experts provide authority through verification on the platform. This division is what allows the system to scale while keeping the final evaluation trustworthy.

Figure 3. Expert grading interface: for each of the six OSATS dimensions the reviewer assigns a score from 1 to 5 and records whether the AI’s justification was a hallucination.
4.5 Outcomes and Quantitative Evaluation Protocol
Over the course of this project, five main deliverables are produced, and a layered quantitative evaluation protocol is established that covers both AI system performance and the reliability of the human ground truth.
4.5.1 Core Deliverables
- Four-level clinical evaluation framework. The perception–workflow–reasoning–scoring framework described in Section 4.3, operationalized as a tiered prompt system for the VLM with a strict structured-output template that pairs every AI prediction with a dedicated human-expert evaluation block.
- Ground-truth templates and expert annotation protocol. Documents that specify exactly how human experts should verify each AI output, including the medical nomenclature to be used and the scoring anchors for each OSATS dimension.
- Medical safety checker. A module that flags potentially dangerous or hallucinated AI recommendations — for example, a predicted next step that contradicts standard surgical protocol.
- Expert grading web platform. The dual-expert independent review platform with automatic disagreement detection and third-party arbitration described in Section 4.4, forming a traceable and reproducible ground-truth pipeline.
- Quantitative evaluation protocol. A protocol covering computer-vision metrics and inter-rater reliability statistics, supporting layered quantitative validation of system performance.
4.5.2 Layered Quantitative Metrics
Evaluation metrics are designed for each of the four framework levels, enabling end-to-end layered quantification.
- Level 1 (visual perception and localization). Operation-type identification is measured by accuracy; anatomical structure and instrument identification are measured by precision, recall, F1, and macro-F1, alongside hallucination rate (objects the model falsely reports as present) and misidentification rate (objects present but named incorrectly); spatial orientation judgment is measured by accuracy.
- Level 2 (sequential workflow modeling). Temporal event localization is measured by temporal Intersection-over-Union (temporal IoU) and mAP@IoU, quantifying the overlap between predicted phase boundaries and ground truth; missed-step detection is measured by precision, recall, F1, and macro-F1.
- Level 3 (multimodal clinical reasoning). Next-step prediction is measured by precision, recall, and macro-F1. Three clinically critical safety metrics are additionally tracked: completion-judgment precision, false-completion rate (the highest-risk clinical error, where an incomplete surgery is judged complete), and global precision, supplemented by terminology accuracy and safety-failure rate.
- Level 4 (OSATS-based skill scoring). Agreement between AI and human OSATS scores is measured by Mean Absolute Error (MAE), Linear Correlation Coefficient (LCC), and Spearman Rank-Order Correlation Coefficient (SROCC). Inter-rater reliability among human experts is measured by Intraclass Correlation Coefficient (ICC) with Fisher z-transformation, supplemented by Kappa tests and Bland–Altman analysis.
All verified outputs are aggregated into a standardized surgical evaluation report: per-level quantitative metrics, the predicted timeline with expert corrections, the completion verdict and next-step prediction, and the six OSATS scores with their visual justifications. These reports can be stored, compared over time, and used for training feedback or quality audits.
5. Preliminary Results
5.1 Dataset Construction Progress
At the time of this report, the curated oral surgery video dataset has completed full source-level enumeration and the first round of instance-level screening. A total of 1,114 videos from 8 distinct sources have been cataloged with complete provenance metadata, including source platform, original URL, publisher, publication year, and licensing status. After applying the standardized inclusion and exclusion criteria to the initial candidate pool, approximately 500 videos have been retained, annotated with standardized procedure nomenclature, and assigned a difficulty level per the established difficulty-gradient classification system. Dataset expansion is ongoing toward the target size of approximately 2,000 videos.
5.2 Formalized Evaluation and Annotation Protocols
The complete quantitative evaluation protocol for the four-level framework has been formalized. Dedicated metrics are defined for each layer: perception accuracy, precision, recall, F1 and hallucination rate for Level 1; temporal IoU and mAP@IoU for Level 2 workflow modeling; next-step precision, false-completion rate and clinical safety indicators for Level 3 reasoning; and MAE, LCC and SROCC for Level 4 OSATS scoring, supplemented by ICC, Kappa and Bland–Altman statistics for inter-rater reliability measurement. In parallel, ground-truth templates and a standardized expert annotation protocol have been finalized, specifying uniform medical nomenclature, OSATS scoring anchors, and step-by-step verification workflows to ensure reproducible expert review.
5.3 Web Platform Pilot Testing
The expert grading web platform has been developed and launched for initial pilot testing. The core functional pipeline has been fully validated: dual independent expert grading, automatic discrepancy detection and highlighting, and third-party arbitration routing all operate as designed. The platform interface integrates dedicated scoring fields for each of the six OSATS dimensions and a dedicated hallucination flag for each AI justification, supporting structured, fully traceable expert verification.
5.4 Case Illustration
Figure 4 demonstrates a complete end-to-end system output on a representative oral surgery video (broken implant abutment screw retrieval, FDI 11). The figure combines input visualization and structured AI output in a single integrated view: the input panel shows the video frame at 00:01:53 (osteotomy preparation phase), overlaid with bounding boxes for detected instruments and anatomy (high-speed handpiece with drill bit, suction tip, implant fixture with gingival cuff), alongside a thumbnail-based temporal segmentation bar mapping seven consecutive phases across the full 00:00–02:56 timeline. The corresponding output panel presents the four-level structured AI results: Level 1 identifies the procedure, 4 anatomical structures, 6 instruments, and spatial orientation; Level 2 reconstructs the 7-phase timestamped timeline and confirms no missed steps; Level 3 judges the surgery incomplete and predicts the correct next action with clinical rationale, with a passed safety check; Level 4 delivers OSATS dimension scores, each grounded in a timestamped visual event. This case validates that the pipeline produces complete, clinically meaningful, expert-verifiable outputs from raw surgical video.

Figure 4 Input Video: The main view is taken from the 01:53 frame (Phase 5: osteotomy preparation phase), overlaid with three detection bounding boxes — high‑speed handpiece + drill bit, suction tip, implant (FDI 11) + gingival cuff. Below lies a thumbnail‑based temporal segmentation bar covering seven phases (00:00→02:56), perfectly aligned with the AI‑generated timeline. Four‑level AI Output: L1 Perception and Localization (surgical procedure, 4 anatomical structures, 6 surgical instruments, spatial orientation); L2 Temporal Workflow (7‑phase timeline + no missing steps); L3 Clinical Reasoning (procedure incomplete + next‑step operation + passed safety check); L4 OSATS‑based Skill Scoring (full score of 5 across 7 dimensions + example timestamp‑supported evidence).
6. Discussion & Future works
This work delivers the first end-to-end multimodal framework for automated structured evaluation of oral surgery videos, paired with a rigorous expert-in-the-loop validation pipeline. In its pilot implementation, the system successfully decomposes raw operative footage into hierarchical clinical outputs — from basic visual perception to workflow reconstruction, clinical reasoning, and standardized skill scoring — and establishes a reproducible pathway to generate trustworthy ground truth through dual expert review and arbitration. Beyond technical performance, the framework addresses a core gap in surgical education and quality assurance by enabling scalable, objective video-based assessment that was previously limited to slow, subjective manual review.
This study exclusively uses the Gemini VLM architecture for three primary reasons. First, Gemini’s native multimodal sequence modeling supports unified tokenization of video frames, audio, and text with temporal positional encoding, which is uniquely suited to the sequential, multi‑sensory nature of surgical video analysis. Second, its long‑context window (approximately 1 M tokens, corresponding to roughly one hour of video) accommodates full‑length procedures without chunking artifacts in the pilot phase. Third, using a single model architecture isolates framework validation as the experimental variable: by holding the base model constant, we can rigorously test the four‑level prompt design, evaluation metrics, and expert verification pipeline without confounding variables introduced by comparing multiple models. Multi‑model comparison is planned for future work now that the framework itself has been validated.
Commercially available general‑purpose multimodal models such as ChatGPT and Doubao are not suitable for this full‑length surgical workflow analysis task, as they lack native long‑duration video ingestion capabilities and rely on sampled frame snapshots rather than continuous temporal video modeling. Similarly, the uAI‑NEXUS‑MedVLM‑1.0a medical VLM exhibits strengths for short‑clip surgical recognition, e.g., direct identification of discrete operative steps from brief video snippets; however, it is less well‑equipped for deep temporal reasoning and holistic procedural understanding across extended operative sequences.
This report itself has several important limitations. The pilot validation is based on only five videos, so statistical power is low and performance results are preliminary and not generalizable to the full dataset or across all procedure types. The evaluation has been conducted exclusively on oral surgery procedures, and generalization to other surgical specialties remains untested. Finally, the dataset and platform have not yet undergone formal multi-institutional clinical validation.
7. Conclusion
This study designed, built, and pilot-validated an end-to-end system that converts unstructured oral surgery videos into standardized, clinically meaningful evaluations. The four-level multimodal AI framework — spanning visual perception, sequential workflow modeling, multimodal clinical reasoning, and OSATS-aligned skill scoring — delivers structured automated assessments, while the expert-reviewed grading platform and embedded safety checker maintain clinical trustworthiness.
The system’s most immediate value lies in surgical education, enabling objective, scalable feedback for trainees and practitioners who would otherwise receive limited, inconsistent review. Future work will expand the evaluation corpus, harden safety guardrails, and test generalization across other surgical specialties, advancing medical AI closer to real-world operating room support.
8. References
[1] Nagda, M., Lee, J., Thompson, M., et al. Towards expert-level medical AI for real-time video consultations. arXiv:2608.09861, 2026.
[2] Martin, J. A., Regehr, G., Reznick, R., et al. Objective structured assessment of technical skill (OSATS) for surgical residents. British Journal of Surgery, 84(2): 273–278, 1997.
[3] Poudel, N., Simon, R., & Linte, Cristian A. (2026). Evaluating Large Vision-language Models for Surgical Tool Detection. arXiv.Org. https://arxiv.org/abs/2601.16895
[4] Nguyen, V. A., Vuong, T. Q. T., & Nguyen, Van H. (2025). Benchmarking large-language-model vision capabilities in oral and maxillofacial anatomy: A cross-sectional study. PLOS One, 20(10), e0335775. https://doi.org/10.1371/journal.pone.0335775
[5] Çetiner, E. Y. (2025). Comparative Evaluation of ChatGPT-5 and Gemini 2.5 Pro in Answering Oral and Maxillofacial Surgery Questions from Dentistry Specialization Exams: A Cross-Sectional Study. EurAsian Journal of Oral and Maxillofacial Surgery, 4(3), 59–65. https://dergipark.org.tr/en/pub/ejoms/article/1775388
[6] Makrygiannakis, M. A., Giannakopoulos, K., & Kaklamanos, E. G. (2024). Evidence-based potential of generative artificial intelligence large language models in orthodontics: a comparative study of ChatGPT, Google Bard, and Microsoft Bing. European Journal of Orthodontics. https://doi.org/10.1093/ejo/cjae017
[7] Aziz, A. A. A., Hams H., A., & Hassan, M. G. (2025). The use of ChatGPT and Google Gemini in responding to orthognathic surgery-related questions: A comparative study. Journal of the World Federation of Orthodontists, 14(1), 20–26. https://doi.org/10.1016/j.ejwf.2024.09.004
[8] Massimi, D., Stefano, L. Di, Rizkala, T., Spadaccini, M., Mori, Y., Menini, M., Antonelli, G., Khalaf, K., Bisschops, R., von Renteln, D., Sharma, P., Rex, D. K., Bretthauer, M., Castoro, C., Repici, A., & Hassan, C. (2026). Large Language Model‐Driven Analysis and Report Generation of Endoscopy Videos—A Pilot Study. Digestive Endoscopy, 38(3). https://doi.org/10.1111/den.70134
[9] Saab, K., Tu, T., Weng, W.-H., Tanno, R., Stutz, D., Wulczyn, E., Zhang, F., Strother, T., Park, C., Vedadi, E., Chaves, J. Z., Hu, S.-Y., Schaekermann, M., Kamath, A., Cheng, Y., Barrett, D. G. T., Cheung, C., Mustafa, B., Palepu, A., … Natarajan, V. (2024, May 1). Capabilities of Gemini Models in Medicine. arXiv.Org. https://doi.org/10.48550/arXiv.2404.18416
[10] Gomez-Cabello, C. A., Sahar Borna, Pressman, S. M., Haider, S. A., & Forte, A. J. (2024). Large Language Models for Intraoperative Decision Support in Plastic Surgery: A Comparison between ChatGPT-4 and Gemini. Medicina, 60(6), 957–957. https://doi.org/10.3390/medicina60060957
[11] Stueker, E. H., Kolbinger, F. R., Saldanha, O. L., Digomann, D., Pistorius, S., Oehme, F., Marko Van Treeck, Ferber, D., Lavinia, M., Weitz, J., Distler, M., Kather, J. N., & Muti, H. S. (2025). Vision-language models for automated video analysis and documentation in laparoscopic surgery: a proof-of-concept study. International Journal of Surgery. https://doi.org/10.1097/js9.0000000000003069
[12] Setzen, S. A., Andreadis, K., Elemento, O., & Rameau, A. (2025). <scp>AI</scp>‐Powered Laryngoscopy: Exploring the Future With Google Gemini. The Laryngoscope, 135(6), 1851–1853. https://doi.org/10.1002/lary.32089
[13] Holste, G., Oikonomou, E. K., Márton Tokodi, Attila Kovács, Wang, Z., & Khera, R. (2025). Complete AI-Enabled Echocardiography Interpretation With Multitask Deep Learning. JAMA. https://doi.org/10.1001/jama.2025.8731
[14] Ban, Y., Eckhoff, J. A., Ward, T. M., Hashimoto, D. A., Meireles, O. R., Rus, D., & Rosman, G. (2024). Concept Graph Neural Networks for Surgical Video Understanding. IEEE Transactions on Medical Imaging, 43(1), 264–274. https://doi.org/10.1109/tmi.2023.3299518
[15] Harari, R. E., Dias, R. D., Kennedy-Metz, L. R., Varni, G., Gombolay, M., Yule, S., Salas, E., & Zenati, M. A. (2024). Deep Learning Analysis of Surgical Video Recordings to Assess Nontechnical Skills. JAMA Network Open, 7(7), e2422520. https://doi.org/10.1001/jamanetworkopen.2024.22520
[16] Chen, Z., Guo, Q., Yeung, L. K. T., Chan, D. T. M., Lei, Z., Liu, H., & Wang, J. (2023). Surgical Video Captioning with Mutual-Modal Concept Alignment. Lecture Notes in Computer Science, 24–34. https://doi.org/10.1007/978-3-031-43996-4_3
[17] Liu, Y., Wang, Z., Li, Y., Liang, X., Liu, L., Wang, L., & Zhou, L. (2024). MRScore: Evaluating Medical Report with LLM-Based Reward System. Lecture Notes in Computer Science, 283–292. https://doi.org/10.1007/978-3-031-72384-1_27
[18] Oguz, C., Hamidullah, Y., Van Genabith, J., & Ostermann, S. (2026). DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding (pp. 38356–38371).https://aclanthology.org/anthology-files/anthology-files/pdf/findings/2026.findings-acl.1912.pdf
[19] Georgenthum, H., Cosentino, C., Marozzo, F., & Liò, P. (2025). Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI. arXiv.Org. https://arxiv.org/abs/2504.19918
[20] Ali, H., Sumon, R. I., Khalid, A. R., Fathima, K., & Kim, H. C. (2025). A Semantic Evaluation Framework for Medical Report Generation Using Large Language Models. Computers, Materials & Continua, 84(3), 5445–5462. https://doi.org/10.32604/cmc.2025.065992
[21] Saeed, T., & Wang, B. (2026). Large language models for generative recommendation: a systematic review of data-centric taxonomy, evaluation, and human-centric analytics. International Journal of Data Science and Analytics, 22(1). https://doi.org/10.1007/s41060-026-01106-9
[22] Neto, J. B. C., Lazarin, R., Matheus, H. R., Santiago, E., & Romito, G. A. (2026). Implantoplasty combined with soft tissue grafting for the management of complex cases: A microsurgical approach. Clinical Advances in Periodontics, 16 Suppl 1(Suppl 1), S97–S110. https://doi.org/10.1002/cap.70037
[23] Abukaraky, A., Hamdan, A.-A., Ameera, M.-N., Nasief, M., & Hassona, Y. (2018). Quality of YouTube TM videos on dental implants. In Medicina oral, patologia oral y cirugia bucal (Vol. 23, Issue 4, pp. e463–e468). PubMed Central. https://doi.org/10.4317/medoral.22447
[24] Schoebrechts, E., de Almeida Mello, J., Vandenbulcke, P., Palmers, E., Declercq, A., Declerck, D., & Duyck, J. (2023). International Delphi Study to Optimize the Oral Health Section in interRAI. Journal of Dental Research, 102(8), 901–908. https://doi.org/10.1177/00220345231156162