Skip to main content
eScholarship
Open Access Publications from the University of California

UCLA

UCLA Previously Published Works bannerUCLA

Detecting Dental Caries Using General-Purpose Large Multimodal Models From Oral Photographs.

Abstract

Objectives

To determine whether a general-purpose large multimodal model (LMM) can detect dental caries from intraoral photographs without task-specific training, evaluating image-level classification and tooth-level localisation under zero-shot (no reference examples) and five-shot (five annotated examples) conditions.

Methods

This diagnostic accuracy study used the benchmark test split of 1255 publicly available intraoral photographs. Gemini 3.1 reasoning and instant were queried via Vertex AI API using stateless calls and structured prompts for sequential, tooth-by-tooth scanning. Each model under different configurations was evaluated across 10 independent inference runs. Predictions (bounding boxes) were evaluated against expert annotations using greedy Intersection-over-Union (IoU ≥ 0.5). True positives required spatial overlap at the tooth level, or at least one correctly localised lesion at the image level. Performance was summarised using sensitivity, precision, F1-score, and mAP@50.

Results

Performance was strongly dependent on image view and model type, with reliable results observed mainly for the reasoning models on occlusal images. At the image level, zero-shot prompting showed high sensitivity but lower precision, yielding values of 0.95, 0.80, and 0.87 for sensitivity, precision, and F1-score, respectively. Five-shot prompting produced a more balanced profile, with corresponding values of 0.88, 0.87, and 0.87. At the tooth level, zero-shot reasoning showed a similar sensitivity-prioritised pattern, with sensitivity, precision, F1-score, and mAP@50 of 0.88, 0.49, 0.63, and 0.62, respectively. Five-shot prompting improved precision and produced a more balanced localisation profile, with corresponding values of 0.77, 0.57, 0.64, and 0.63.

Conclusions

General-purpose LMMs such as Gemini show potential as a scalable, automated caries screening tool for teledentistry without domain-specific fine-tuning. Although prompt engineering effectively modulates the sensitivity-precision trade-off, the model's tendency toward over-detection requires rigorous clinical validation and governance before real-world deployment as a public screening tool.

Many UC-authored scholarly publications are freely available on this site because of the UC's open access policies. Let us know how this access is important for you.