Abstract
Aims and background: Artificial intelligence (AI)-assisted clinical documentation systems have demonstrated efficiency and structural benefits in medical settings; however, empirical evidence regarding their performance in dental clinical documentation remains limited. Dentistry requires highly granular procedural and anatomical detail, raising concerns about the reliability and adequacy of AI-generated notes. This study compared the accuracy, completeness, clarity, and specificity of AI-generated dental documentation with provider-generated notes under standardised conditions and evaluated statistical reliability, efficiency, and clinician perceptions. Materials and methods: A controlled comparative study design was implemented using standardised scripted audio recordings representing common outpatient dental encounters. Ten licensed dentists documented a denture follow-up encounter based solely on audio recall without note-taking. The identical encounter transcript was processed using an AI-based clinical documentation system. Documentation quality was independently evaluated using a structured four-domain rubric assessing accuracy, completeness, clarity, and specificity on a five-point ordinal scale. Non-parametric statistical analyses were conducted following normality testing, and inter-rater reliability was assessed using intraclass correlation coefficients (ICCs). Qualitative clinician feedback was analysed thematically. Results: Composite documentation scores deviated significantly from normality (Kolmogorov–Smirnov p ≤ 0.001; Shapiro–Wilk p ≤ 0.004). Wilcoxon signed-rank analysis demonstrated a statistically significant overall advantage for AI-generated documentation (Z = −3.744, p < 0.001). Across 40 paired domain comparisons, AI notes were rated higher in 23 instances, provider notes in two instances, and 15 comparisons were tied. At the encounter level (n = 10), AI was rated higher in seven encounters for accuracy, completeness, and clarity, and in six encounters for specificity. A strong positive correlation was observed between provider and AI composite scores (Spearman ρ = 0.756, p < 0.001). Post-hoc power analysis demonstrated high achieved power (0.988). Qualitative analysis identified recurrent omissions in dentistry-specific details and perceived excessive verbosity in AI-generated notes. Conclusion: Artificial intelligence-generated dental documentation can match or exceed provider-generated notes on structured quality metrics under controlled conditions. However, dentistry-specific procedural granularity remains inconsistently captured. Artificial intelligence systems should therefore be implemented within supervised, human-in-the-loop workflows to preserve clinical precision and medico-legal defensibility. Clinical significance: Artificial intelligence-assisted documentation systems may improve the efficiency and quality of dental record keeping. However, clinician verification remains necessary to ensure the accurate capture of dentistry-specific findings and maintain medico-legal standards.