Abstract
Objective
This study aims to develop a personalized automated system for annotating three-dimensional (3D) facial soft tissue landmarks utilizing deep learning and computer vision. By comparing the results of automated annotations with manual annotations, the accuracy and clinical applicability of the proposed algorithm were systematically evaluated, providing an efficient and precise analysis system for facial morphology research.
Methods
A total of 55 Chinese orthodontic patients (24 males, 31 females; mean age 23.4 ± 7.01 years) were recruited from Shandong University Stomatological Hospital, comprising 40 patients with normal facial morphology and 15 with severe craniofacial deformities (mandibular asymmetry, severe skeletal Class II and III; n = 5 each). This study constructed a personalized automated system based on deep learning and computer vision. Through standardized facial template construction with 68 key points, automated 68-landmark annotation of original scans, 3D facial nonlinear registration, and personalized keypoint transfer, the system enables one-time personalized annotation of the standard template and subsequent automatic batch mapping to multiple facial scan models. Deformity-specific personalized templates were constructed for severe malocclusions. Annotation accuracy was evaluated against manual annotations by experienced orthodontic experts using 21 clinically relevant landmarks.
Results
The system demonstrated outstanding accuracy with an average Euclidean distance of 0.95 ± 0.48 mm for 21 selected landmarks, surpassing most existing algorithms. Proportion-based analysis revealed that 96.3% of all landmarks achieved clinically acceptable accuracy within 2.0 mm, with 84.0% within 1.5 mm and 57.0% achieving sub-millimeter precision. The average errors for linear and angular measurements were 0.31 ± 0.35 mm and 1.05 ± 0.90°, respectively. For mandibular deformity samples, personalized templates significantly improved accuracy, reducing mean Euclidean distance from 0.95 ± 0.33 mm (standardized template) to 0.84 ± 0.21 mm (personalized template), with significant error reductions across all 21 landmarks, particularly in mandibular regions.
Conclusions
The proposed automated annotation system achieves efficient and precise annotation of 3D facial soft tissue landmarks through standardized facial template construction, 3D facial nonlinear registration, and personalized landmark transfer. This system provided significant convenience for the measurement of three-dimensional landmarks of soft tissues, the analysis of facial soft tissue morphology, and the diagnosis, treatment, and efficacy evaluation of malocclusion. Additionally, The personalized template approach enhances annotation precision for severe facial deformities, broadening clinical applicability.
Similar content being viewed by others
Introduction
With the increasing emphasis on both physical and mental health, coupled with the shift towards a biopsychosocial medical model, modern orthodontic treatment has expanded its focus from occlusion to facial soft tissue aesthetics. This aesthetic-oriented orthodontic philosophy has driven the rapid development of facial morphology measurement technologies [1, 2].
In facial morphology research, the accurate localization of soft tissue landmarks is crucial for effective measurement and analysis, particularly in preoperative orthodontic diagnostics, target position prediction, and postoperative comparison studies [1,2,3,4,5]. Among these, the 3dMD stereophotogrammetry system stands out for its ability to capture 360° facial images in a short time, making it widely used in fields such as facial aging research, attractiveness evaluation, and orthodontic treatment planning [3,4,5]. It is considered the "gold standard" for stereophotogrammetry, providing facial models highly consistent with real facial morphology for reliable craniofacial morphological studies [6,7,8].
In the study of facial morphology, anatomical landmarks of the face and their relative positional relationships are crucial for describing the morphological characteristics of oral and maxillofacial soft tissues and determining relevant measurement parameters. These landmarks serve as the premise and foundation for three-dimensional facial morphological analysis. Liu et al. [9] evaluated the three-dimensional morphological changes in the vermilion zone of patients after orthodontic treatment, both with and without tooth extraction, based on perioral anatomical landmarks and the linear distances and surface areas they form. Lambros et al. [10] explored methods for quantitatively describing facial morphological changes during aging, such as the appearance of wrinkles and the deepening of tear troughs, based on periorbital and midfacial anatomical landmarks. Additionally, Du et al. [11] investigated the relationship between the distinctive facial morphology of acromegaly patients and their disease state through linear and angular measurement indices formed by facial landmarks, aiming to guide treatment planning. Therefore, the accurate localization of soft tissue landmarks is of great significance for maxillofacial morphological analysis, pre-treatment diagnosis, treatment planning, prediction of target positions post-treatment, and efficacy evaluation.
However, traditional landmark localization methods, which primarily rely on manual operations by practitioners, are time-consuming and labor-intensive, subject to subjective variability, and lack consistency and repeatability. With the increasing prevalence of 3D facial morphology analysis, there is an urgent need for efficient, accurate, and robust automated landmark localization methods. Traditional geometric algorithms and model-matching approaches have shown limited flexibility and adaptability, especially when dealing with complex facial structures.
Deep learning technologies, with their strong feature extraction capabilities, have recently shown significant advantages in medical image analysis [12, 13]. Automated annotation technologies based on deep learning have matured in cephalometric radiography and CBCT landmark recognition, outperforming traditional algorithms in efficiency and accuracy. However, research on 3D facial soft tissue landmarks remains in its infancy. For example, Bo Berends et al. [14] used the DiffusionNet algorithm to automatically annotate 10 soft tissue landmarks on 3D facial photographs and subsequently evaluated its accuracy. Yang et al. [15] employed a high-resolution network to automatically detect and validate perioral landmarks. These studies highlight the immense potential of deep learning methods for improving annotation precision and efficiency. However, most existing studies focus on specific landmark localization under fixed templates, without adequately addressing the anatomical variations and personalized needs of different patients.
To address these issues, this study aims to develop a personalized automated annotation system for 3D facial soft tissue landmarks based on a hybrid framework integrating deep learning-based landmark detection and optimization-based computer vision techniques. Using 3dMD stereophotogrammetry to capture 3D facial images, we constructed a standardized and personalized landmark template and developed an efficient non-rigid registration algorithm to meet the personalized annotation needs of patients with diverse anatomical structures. By comparing automated and manual annotation results, this study systematically evaluated the proposed algorithm's accuracy and clinical applicability, offering efficient and precise analysis system for orthodontic treatment and facial morphology research.
Materials and methods
Acquisition of facial scan data
Study population and ethical approval
This retrospective study collected three-dimensional (3D) facial image data from 55 Chinese orthodontic patients who visited the Orthodontics Department of Shandong University Stomatological Hospital between July 2024 and October 2025, including 24 males and 31 females with a mean age of 23.4 ± 7.01 years. Patients with a history of craniofacial trauma or cleft lip/palate were excluded from the study. This research received approval from the Ethics Committee of Shandong University (approval number: 20241006) and adheres strictly to the ethical principles for human research outlined in the Declaration of Helsinki (2013 revised edition).
The sample size for the normal facial morphology group was determined through a priori power analysis using PASS 15.0 software (NCSS, LLC, Kaysville, Utah, USA). Based on pilot study data showing a mean difference of 1.04 mm between automated and manual landmark annotations (pooled SD = 0.41 mm), and using an equivalence margin of 1.3 mm, a paired equivalence t-test was designed with α = 0.05 (two-sided) and power of 1-β = 0.90. The analysis indicated that a minimum of 38 subjects would be required. To account for potential data loss, 40 subjects with normal facial morphology were recruited, which provided 91.66% statistical power as confirmed by post-hoc analysis. Additionally, 15 subjects with severe facial deformities were included for preliminary exploratory evaluation of the personalized template approach (Sect. 3.5), bringing the total sample size to 55 cases.
Sample stratification and grouping
The study population (n = 55) was stratified into two primary groups based on facial morphological characteristics and malocclusion severity (Table 1).
-
(1)
Standardized Template Evaluation Group (n = 40): Patients with no apparent facial deformities or severe malocclusions, including those with balanced facial profiles and minor malocclusions not requiring surgical intervention. This group was used to evaluate the accuracy of the standardized template-based automated annotation system.
-
(2)
Personalized Template Evaluation Groups (n = 15): Patients with clinically significant facial deformities, including mandibular asymmetry (menton deviation > 3 mm from midline [16]), severe skeletal Class II (ANB > 6°) [17], or severe skeletal Class III (ANB < −4°) [18]. All 15 cases were used to evaluate the accuracy of three pre-constructed personalized templates corresponding to these three deformity types.
3D Facial image acquisition protocol
During the scanning process, participants were instructed to pull their hair back and remove any eyeglasses, and jewelry. Subjects sat upright in a postural position with relaxed lips and forward-facing gaze. Facial images were captured using the 3dMD stereo system within a standardized room environment. The intensity and angle of the light source were kept consistent, and the collection environment was designed to minimize factors that could cause reflections. The facial data was subsequently preprocessed (including denoising and optimization) and optimized using 3dMD software, then exported in Object Wavefront format ('.obj' file) to ensure consistency and efficiency in further analysis.
Manual landmark placement
The preprocessed 3D images were imported into 3DSlicer version 5.6.2 in Obj format. Two orthodontic experts, each with over five years of experience and systematic training, independently performed the manual annotation of 21 commonly used soft tissue landmark points in each image. The definition of the landmarks adheres to international standards (Table 2, Fig. 1), and the x, y, and z coordinates of each landmark point were recorded.
Manual landmark placement on slicer5.6.2 software
To ensure reliability of the reference annotations, two orthodontic experts independently labeled all landmarks. Any discrepancies between the two raters were resolved through consensus discussion. In cases where consensus could not be reached, a senior orthodontic expert with over ten years of clinical experience and a higher professional rank adjudicated the final landmark position. This process ensured that the final manual annotations represented the most accurate and reproducible reference standard for subsequent model evaluation.
Automated landmark placement
The automated landmark process comprises several steps, which are illustrated in Fig. 2.
Algorithm development workflow. A Acquisition of Facial Scan Data: High-precision facial scanning data is obtained through the 3dMD facial scanning device to generate a 3D facial model. B Key points Annotation of Original Facial Scans: the VTK library is employed to produce a two-dimensional virtual frontal image, while 68 key points are automatically annotated using the Dlib algorithm. b The projection of these two-dimensional key points onto the surface of the three-dimensional facial model through the ray casting algorithm. C Standardized face template construction and registration: Select BFM as the standardized template to perform the registration process. c Details of the registration process: The registration process is divided into two stages rigid alignment and nonlinear registration. Initially, the 68 key points are utilized for rigid alignment, resulting in initial registration outcomes. Subsequently, a network featuring a double-loop optimization structure is implemented for nonlinear registration. This stage incorporates linear registration techniques by introducing vertex distance, key point distance, stiffness loss, and Laplacian smoothing loss, thereby ensuring global shape matching while preserving local details, as well as maintaining registration accuracy and grid smoothness. Nonlinear registration inner loop: optimize the top of the local region shape, and generate optimized vertices based on the local radiation transformation formula. D Registration error analysis and landmark points transfer: utilizing the registration matrix produced through non-rigid registration to transfer personalized landmarks from the standardized template to the original face, with the accuracy of the landmark point transfer evaluated based on the registration error heat map.
Construction of a standardized facial template comprising 68 key points
Standardized face templates encompass the Basel Face Model (BFM) and personalized templates. The BFM model features 68 key points [19], offering high accuracy, and is extensively utilized in the domains of facial modeling and recognition [20, 21].
The personalized template is constructed based on step 3.5, during which an experienced dental expert utilizes 3DSlicer software to accurately adjust the positions of 68 key points on the template. Following this, two dental experts independently review the adjustments to ensure the accuracy and consistency of the annotations.
68-Landmark annotation of original facial scans
In the rigid alignment stage outlined in step 3.2, the 68 key points of the original face are utilized. This process involves employing the Visualization Toolkit (VTK) library to conduct virtual frontal photography on the preprocessed facial scan data. The resulting two-dimensional frontal image is annotated with high precision using the Dlib library, which automatically identifies the 68 key points. Following this, the ray casting method is applied to accurately project these key points from the two-dimensional image onto the three-dimensional facial surface, thereby determining their positions in three-dimensional space.
Optimization-based non-rigid local affine registration
This system employs a deep-learning-based non-rigid local affine registration model [22] for nonlinear registration. The double-layer cyclic structure of the registration model facilitates global shape registration while preserving the accuracy of local details.
It is important to note that although PyTorch is utilized for implementation, the registration component employs optimization-based methods with automatic differentiation rather than neural network training. The deep learning aspect of our system is specifically applied in the pre-trained Dlib model for 2D landmark detection.
Rigid alignment
The global rigid alignment was achieved by optimizing a rotation matrix and a translation matrix to minimize vertex errors between the standard and original facial meshes, ensuring an initial alignment based on the 68 landmarks.
The optimization problem can be expressed as:
Where:
\(\left({p}_{i}\right)\) is the (i) -th vertex (key point) of the source mesh (standard face).
\(\left({q}_{i}\right)\) is the (i)-th vertex (key point) of the target mesh (original face).
\(\left(N\right)\) is the number of corresponding key points (in this case, (N = 68)).
Non-rigid transformation optimization
Utilizing the local affine transformation model, along with the affine transformation matrix and translation vector, the optimization of facial detail deformation is performed hierarchically. Each vertex is adjusted flexibly while preserving the local structural characteristics.
-
(1)
Outer Loop: The model iteratively refined overall shape matching through N phases, gradually reducing shape discrepancies between source and target meshes.
-
(2)
Inner Loop: Each outer loop stage involved multiple inner loop optimizations, adjusting vertex positions in local regions to better align with the target mesh.
Where:
\(\left({p}_{j}\right)\) is the (j)-th vertex of the of the local region.
\(\left(A\right)\) is the local affine transformation matrix for the (j)-th vertex.
\(\left(B\right)\) is the local translation vector for the (j)-th vertex.
\(\left({q}_{j}\right)\) is the (j)-th vertex of the of the local neighbors.
\(\left(L\right)\) is the number of vertices in the local region.
Loss function design
The registration process was optimized using multiple loss functions (3),
including:
-
(1)
Vertex Distance Loss \(\left({L}_{vertex}\right)\): Governs overall shape alignment (4);
Where:
\(\left({p}_{j}\right)\) is the (j)-th vertex of the source mesh.
\(\left({q}_{j}\right)\) is the (j)-th vertex of the of the target mesh.
(M) is the number of vertices of the of the source mesh.
-
(2)
Landmark Distance Loss \(\left({L}_{landmark}\right)\): Ensures precise landmark alignment (5);
Where:
\(\left({p}_{i}\right)\) is the (i)-th transformed vertex (keypoint) of the source mesh.
\(\left({q}_{i}\right)\) is the (i)-th vertex (keypoint) of the target mesh.
\(\left({w}_{i}\right)\) is the weight for each pair of key points.
\(\left(N\right)\) is the number of corresponding key points.
-
(3)
Stiffness Loss \(\left({L}_{stiffness}\right)\): Maintains local rigidity, preventing unrealistic deformation (6);
Where:
\(\left({A}_{k}\right)\) is the local affine transformation matrix of the (k)-th region.
(I) is the identity matrix.
\(\left({w}_{k}\right)\) Is the weight of the rigidity loss of the (k)-th region.
\(\left(K\right)\) is the number of local regions.
-
(4)
Laplacian Smoothing Loss \(\left({L}_{lap}\right)\): Enhances mesh smoothness, avoiding unnatural mesh deformation (7).
Where:
\(\left({p}_{i}\right)\) is the (i)-th vertex of the source mesh.
\(\left({p}_{j}\right)\) is the (j)-th vertex in the neighboring vertex set \(\left({U}_{(i)}\right)\) of pi.
\(\left(\left|{U}_{\left(i\right)}\right|\right)\) is the number of neighboring vertices of \({p}_{i}\).
(M) is the number of vertices of the of the source mesh.
Hyperparameter configurations
Our double-loop registration method employs three hyperparameter configurations tailored to different application scenarios. Table 3 summarizes the iteration counts, weight scheduling strategies for each configuration.
Transfer of personalized landmarks
Upon registration completion, we transferred the 21 soft tissue landmark points of the individualized template to the original facial model using the computed transformation matrix, thereby generating personalized landmark point annotations. The accuracy of these annotations is assessed through error heatmaps and vertex error files. In instances where the error surpasses the established threshold (mean value < 0.15 mm, maximum value < 0.20 mm), the registration process is repeated by modifying the registration parameters.
Selection and fine adjustment of personalized templates (optional)
For facial models with significant anatomical discrepancies such as pronounced mandibular retrusion or facial asymmetry, a sample can be selected from this category to serve as a personalized template, thereby ensuring the accuracy and applicability of personalized landmark point annotation.
In this study, three deformity-specific personalized templates were independently constructed using representative cases not included in the evaluation dataset: (1) a mandibular asymmetry template, (2) a severe Class II template, and (3) a severe Class III template. For each template, experienced orthodontic experts meticulously adjusted the 68 key points, and constructed it as a personalized template using the method described in Sect. 3.1. We marked the 21 landmark points on this template.
These template-creation cases were strictly separated from the 15 cases in the evaluation groups. This separation ensures no data leakage occurred between template construction and performance validation, allowing for an unbiased assessment of the personalized template approach.
Implementation details
Registration model architecture
The model consists of learnable transformation parameters without neural network layers:
-
Aᵢ ▯ℝ3ˣ3: Local affine matrix (initialized as identity)
-
bᵢ▯ℝ3ˣ1: Translation vector (initialized as zero)
-
Total: 53,490 vertices × 12 parameters = 641,880 parameters
-
Forward transformation (Eq. 2 in manuscript): min Σ |qⱼ—(Apⱼ + B)|2
Template meshes (source) composition
The registration dataset comprises template meshes and target scans.We employ two template types depending on patient anatomy:
-
(1)
Standardized Template: The Basel Face Model (BFM) [19] serves as the default template, featuring 53,490 vertices, 106,466 faces, and 68 manually annotated landmarks. This template is suitable for patients without significant anatomical abnormalities.
-
(2)
Personalized Template: For severe facial deformities or specific malocclusion types, we construct patient-specific templates from user-annotated models. These templates are stored as.obj files (arbitrary topology) with landmark indices in.npy format.
Optimization algorithm and convergence criteria
We employ the AdamW optimizer with a learning rate of 1 × 10⁻4 and AMSGrad enabled for stable convergence [23]. Registration quality is assessed using vertex-to-surface distances; results are accepted if mean error < 0.15 mm and maximum error < 0.20 mm, otherwise flagged for manual review.
Computational environment
-
(1)
Hardware and Software: The system was deployed on a workstation equipped with dual Intel Xeon Gold 5218R processors (40 cores), NVIDIA A100-40 GB GPU, and 128 GB RAM. The implementation utilized Python 3.8.10 with PyTorch 1.12.0, CUDA 11.6, Dlib 19.22.0 (pre-trained 2D landmark detection).
-
(2)
Performance: Average inference time was 30–45 s per case with approximately 2.5 GB GPU memory consumption.
-
(3)
Reproducibility: The optimization-based registration is deterministic with fixed hyperparameters, requiring no random seed initialization.
Anthropometric measurement
After all landmark points have been marked, their 3D coordinates are exported to Microsoft Excel 2022 (Microsoft Corporation, Redmond, WA). Subsequently, five linear distances and five angular measurements [15, 24] based on these landmark points are calculated to further evaluate the accuracy of the system(Fig. 3).
Linear and angular measurements are shown on a 3D facial photograph in frontal view (A) and in lateral view (B). Five angular measurements: 1. lip spread angle: Cheilion-left to Subnasale to Cheilion-right. 2. Central bow angle: Crista Philtra-left to Labiale Superius to Crista Philtra-right. 3. Nasolabial angle: Pronasale to Subnasale to Labiale Superius. 4. Labiale Superius to Stomion to Labiale inferius. 5. Labiale inferius to Sublabiale to Soft tissue Pogonion. Five linear measurements:6. Alare width: Alare-left to Alare-right. 7. Philtrum width: Sublabiale to Labiale Superius. 8. Lip width: Cheilion-left to Cheilion-right. 9. Upper vermilion height: Subnasale to Labiale Superius. 10. Labiale Superius to Stomion
Statistical analysis
Data analysis was conducted using Microsoft Excel (version 2022) and SPSS (version 26), and statistics were presented as mean ± standard deviation. Statistical assessments were performed on the Euclidean distance between manual and automatic annotation results, with the mean error and standard deviation calculated for each landmark point. To enhance clinical interpretability, For each landmark, we determined the percentage of the 40 test cases where the Euclidean distance between automated and manual annotations fell within each threshold range(≤ 0.5 mm, ≤ 1.0 mm, ≤ 1.5 mm, and ≤ 2.0 mm). Additionally, the differences in coordinate values (x-axis, y-axis, and z-axis) and anthropometric measurements (linear and angular) derived from the landmark points were analyzed to provide a comprehensive evaluation of the algorithm's accuracy and stability.
Prior to comparative analyses, the normality of data distribution was assessed using the Shapiro–Wilk test with a significance level of α = 0.05. For normally distributed data, paired t-tests were employed to compare the coordinate differences and measurement values between manual and automated annotations. For non-normally distributed data, two-tailed Wilcoxon signed-rank tests were applied. The choice between parametric and non-parametric tests for each variable is indicated in the results tables with superscript annotations (a for paired t-test, b for Wilcoxon signed-rank test). Intra-observer and inter-observer reliability were evaluated using intra-class correlation coefficients (ICC) with a two-way random effects model and absolute agreement definition. For all statistical tests, the significance level (α) was set at 0.05, with P < 0.05 considered statistically significant. All statistical tests were two-tailed unless otherwise specified.
Comparative analysis with state-of-the-art methods
To validate the performance of our proposed method against existing state-of-the-art approaches, we conducted direct comparisons with two recently published deep learning methods for automated 3D facial landmark detection: DiffusionNet [14] and 2S-SGCN [25]. We evaluated all methods on identical test sets to ensure fair comparison.
Sample size for comparative evaluation with DiffusionNet and 2S-SGCN was determined through a priori power analysis using PASS 15.0 software. Based on pilot data showing annotation errors of 1.67 ± 1.09 mm (DiffusionNet) versus 1.12 ± 0.74 mm (our method), we calculated that 59 cases per group would provide 90% power to detect this difference at α = 0.05 (two-sided) in a paired design. The Headscape dataset comparison included 60 cases to meet this requirement.
Comparison with DiffusionNet
We first compared our method with DiffusionNet, a recent diffusion-based deep learning approach for 3D facial landmark detection. To ensure a fair comparison, we utilized the official implementation and pre-trained model parameters provided by the original authors. Both methods were evaluated on two independent datasets: 60 3D facial photographs randomly selected from the Headspace database and 40 photographs from our proprietary dataset.
Due to architectural constraints, DiffusionNet can predict 10 soft tissue landmarks: bilateral exocanthions, endocanthions, nasion, nose tip, alares, and cheilions. The Euclidean distances between automated and manual annotations were calculated for both methods. To account for multiple comparisons across the 10 landmarks, we applied Bonferroni correction to control the family-wise error rate, setting the adjusted significance level at α = 0.005 (0.05/10). Paired t-tests were performed to compare the prediction errors between the two methods for each individual landmark.
Comparison with 2S-SGCN
We further compared our method with 2S-SGCN [25], a two-stage spatial graph convolutional network approach for dense 3D facial landmark detection. Since the 2S-SGCN authors also evaluated their method on the Headspace dataset, we compared the mean Euclidean distance achieved by our method (n = 60) directly against the performance metrics as reported in their publication This comparison utilized the 68 overlapping landmarks to assess global performance.
Results
Error analysis
The mean “vertex-to-surface distance” of the nonlinear registration between standardized and original faces was less than 0.2 mm (Fig. 4), indicating a high-fidelity mesh alignment.
Nonlinear alignment error between standardized and original faces
The intra-observer ICC values for each landmark ranged from 0.775 to 0.972, while the inter-observer ICC values ranged from 0.756 to 0.976, confirming overall good reliability of the measurements. However, notable variability was observed across different landmarks, with periocular landmarks (Endocanthion and Exocanthion) demonstrating the lowest ICC values (0.756–0.878), particularly along the Y-axis. In contrast, midline landmarks such as Nasion, Pronasale, and Stomion exhibited consistently high reliability across all axes (Table 4).
Standard template accuracy
For 40 individuals with no apparent facial deformities, the Euclidean distances between manual and automated annotations were calculated for 21 landmarks to assess the accuracy (Table 5 and Fig. 5). The average Euclidean distance for these 21 selected landmark points was 0.95 ± 0.48 mm. Notably, the Euclidean distance between the automated and manual annotation of the bilateral Endocanthion and Exocanthion was the only measurement that exceeded 1 mm.
Euclidean distances between manual and automated landmarks using the standardized template
Proportion-based analysis revealed that the automated system achieved high clinical accuracy across most landmarks(Table 6). Overall, 96.3% of the landmarks were predicted within 2.0 mm. Among the 21 landmarks, Nasion of Soft tissue demonstrated the highest precision, with 100% within 1.5 mm. In contrast, the bilateral Endocanthion and Exocanthion landmarks exhibited relatively lower precision, with only 15–25% of predictions within 1.0 mm. However, even for these challenging landmarks, 70–77.5% of predictions remained within the clinically acceptable 2.0 mm threshold.
We evaluated the differences between manually marked landmark points and those identified by this automated algorithm across the three dimensions: x-axis, y-axis, and z-axis (Table 7). The average error in the x-axis and z-axis direction was 0.54 ± 0.42 mm and 0.39 ± 0.34 mm, respectively, with no statistically significant differences observed between all marking points of the automated algorithm and manual marking. In contrast, the average error on the y-axis was 0.63 ± 0.50 mm; notably, statistical differences existed only for the Endocanthion (Left) and Exocanthion (Right) points, while no significant differences were found for the other marked landmarks.
Additionally, we reported the differences in linear and angular measurements between selected landmark points after manual and automated marking to further evaluate the accuracy of these annotation methods (Table 8). The average prediction errors for linear measurements and angular measurements were 0.31 ± 0.35 mm and 1.05 ± 0.90°. Across all measurement metrics, no statistically significant differences were observed.
Comparative performance with diffusionnet and 2S-SGCN
Tables 9 and 10 present the comparative performance of our method and DiffusionNet on 10 overlapping soft tissue landmarks across two independent datasets. On the Headspace database (n = 60), our method achieved a mean Euclidean distance of 1.08 ± 0.63 mm compared to 1.65 ± 1.09 mm for DiffusionNet. The most pronounced improvements were observed for midline and lower facial landmarks, including nasion, nose tip, bilateral alares, and bilateral cheilions. For canthal landmarks (endocanthions and exocanthions), although the differences between the two methods did not reach statistical significance after Bonferroni correction in most comparisons, our method still demonstrated reduced mean errors. Similar performance patterns were observed on our proprietary dataset (n = 40), with our method achieving 1.07 ± 0.54 mm compared to 1.53 ± 0.87 mm for DiffusionNet (p = 0.003).
To further validate our method's performance across a comprehensive set of facial landmarks, we compared it with 2S-SGCN [25]. On Headspace dataset (n = 60), our method achieved a mean Euclidean distance of 1.02 ± 0.68 mm across all 68 landmarks, compared to 1.66 ± 0.73 mm for 2S-SGCN in overall accuracy. This comparison validates that our method maintains high accuracy even when predicting a dense set of landmarks distributed across the entire facial surface.
Personalized template accuracy
To evaluate personalized template performance across severe craniofacial deformities, we compared annotation accuracy between standardized BFM and deformity-specific personalized templates in 15 cases using a paired design (Table 11). Personalized templates significantly improved overall annotation accuracy, reducing mean Euclidean distance from 0.95 ± 0.33 mm to 0.84 ± 0.21 mm.
Subgroup-specific analyses revealed consistent improvements across deformity types: mandibular asymmetry cases demonstrated 12.6% error reduction, with greatest gains in mandibular landmarks including Soft tissue Pogonion, Menton, and Sublabiale; severe Class II cases showed 11.0% improvement, particularly in perioral and mandibular regions including Soft tissue Pogonion, Labiale inferius, and Labiale Superius; severe Class III cases exhibited 11.2% enhancement, with pronounced improvements in prognathic mandibular features including Soft tissue Menton, Sublabiale, and Gnathion.
Discussion
In this study, we developed a personalized automated annotation system for three-dimensional (3D) facial soft tissue landmarks based on deep learning and computer vision. This system efficiently achieved personalized landmark annotation by combining a standardized facial template with a nonlinear registration algorithm. This approach eliminated the need for extensive manual annotation, significantly reducing labor costs and improving efficiency. The system also addressed the challenges of complexity in existing methods, achieving performance comparable to that of experienced clinicians' manual annotations, exhibiting minimal error, and demonstrating considerable potential for a wide range of clinical applications.
Most existing literature indicated that craniofacial landmarks localization errors within 2 mm are clinically acceptable [26,27,28,29], and Hajeer et al. have suggested that a standard deviation of ≤ 0.5 mm in landmarks localization error signifies high repeatability [30, 31]. The experimental results of this study demonstrated remarkable precision: the mean error in the 3D coordinates of 21 landmarks was 0.5 ± 0.47 mm, and the Euclidean distance was 0.95 ± 0.48 mm, significantly outperforming most current algorithms (1.4–3.96 mm). Notably, landmarks such as Nasion of Soft Tissue, Pronasale, and Subspinale exhibited errors of less than 0.4 mm. Proportion-based analysis further validated clinical utility, with 96.3% of all landmarks within the 2.0 mm threshold, 84.0% within 1.5 mm, and 57.0% achieving < 1.0 mm precision—results that compare favorably with those of Berends et al. [14] and approach the inter-operator variability of manual annotation reported by Fagertun et al. [29] (50–60% within 1.0 mm). This proportion-based metric facilitates informed decision-making regarding the system's application in various clinical scenarios: landmarks with > 70% precision within 1.0 mm can be confidently used for detailed morphometric analyses, while those within 1.5–2.0 mm remain suitable for broader diagnostic and treatment planning purposes. Regarding the three-dimensional changes in the coordinate values of landmark points, it was noteworthy that only the Endocanthion (Left) and Exocanthion (Right) exhibit statistically significant differences on the Y-axis. The remaining landmark points, as well as the measurements of linear and angular based on these landmarks, did not show significant differences when compared to the manual labeling method. These findings indicated the reliability and stability of the system in landmark annotation. However, the precision of Endocanthion and Exocanthion remained an area for improvement, where the error was significantly greater than that of other landmark points (Euclidean distance > 1 mm), particularly along the Y-axis. ICC analysis revealed a direct correlation between manual annotation reliability and automated system performance for these periocular landmarks: Endocanthion and Exocanthion demonstrated substantially lower ICC values compared to other facial features, with Y-axis reliability being particularly compromised. Critically, these were the same landmarks that exhibited the highest prediction errors in our automated system, suggesting inherent challenges in defining precise reference standards even during manual annotation by experienced clinicians. This concordance between reduced manual reliability and increased algorithmic error indicates that the automated system's performance for periocular landmarks reflects, in part, the anatomical ambiguity of these features rather than solely algorithmic limitations. This observation aligned with the findings of Galvánek et al. [32]. The higher errors in these regions may be attributed to their anatomical characteristics, where Endocanthion and Exocanthion converge at the junction of upper and lower eyelid margins. This region contains intricate anatomical structures, such as the medial palpebral ligament, and exhibits variability across individuals, likely contributing to the observed discrepancies.
Mainstream algorithms, such as DiffusionNet, often rely on extensive training with annotated datasets to achieve automatic landmark localization [33,34,35,36]. However, their applicability is constrained in scenarios with limited data or complex anatomical features, and they are unable to identify landmarks not included during training. To further validate the clinical applicability of our proposed system, we conducted direct comparisons with two state-of-the-art deep learning methods: DiffusionNet and 2S-SGCN. Our method demonstrated statistically significant superior performance compared to DiffusionNet across both the Headspace dataset and our proprietary dataset, with the most pronounced improvements observed for midline and lower facial landmarks such as nasion, nose tip, bilateral alares, and bilateral cheilions. These landmarks typically exhibit more distinct geometric features and consistent anatomical positioning, which may be better captured by our geometry-based registration approach combined with personalized template matching. When compared with 2S-SGCN across 68 comprehensive facial landmarks, our method achieved a mean Euclidean distance of 1.02 ± 0.68 mm versus 1.66 ± 0.73 mm, demonstrating consistent superiority across dense landmark sets. These comparative results underscore several key advantages of our approach. First, unlike purely data-driven deep learning methods requiring extensive training datasets, our registration-based framework achieves high accuracy with minimal personalized annotation effort—requiring only a single annotation of the standard template. Second, the modular design allows for flexible adaptation to diverse clinical scenarios and anatomical variations through personalized templates, a feature that fixed-architecture neural networks cannot easily accommodate. Third, our method maintains interpretability and traceability throughout the annotation process, with each registration step being independently verifiable, which is crucial for clinical validation and quality control. However, for canthal landmarks, while our method showed reduced mean errors compared to DiffusionNet, the differences did not consistently reach statistical significance after Bonferroni correction, aligning with the inherently challenging nature of canthal region localization due to complex anatomical structures and high inter-individual variability affecting all automated methods to varying degrees. Moreover, our system provides significant advantages in expanding the number of landmarks, accommodating additional landmarks such as Alare, Subspinale and Sublabiale, ensuring broader applicability across diverse clinical scenarios while maintaining high localization precision for newly added landmarks.
Furthermore, the introduction of personalized templates allows the system to adapt to the specific anatomical characteristics of individual patients. Experimental results indicated that deformity-specific personalized templates significantly improved annotation accuracy across severe craniofacial deformities, with greatest improvements in anatomical regions most affected by each deformity type. This anatomically specific performance pattern validates that personalized templates accommodate morphological deviations from population norms by encoding characteristic deformity features, thereby resolving systematic mismatches encountered with standardized templates. Once constructed through one-time manual annotation, deformity-specific templates enable efficient batch processing of similar cases, offering particular value for specialized orthognathic centers treating high volumes of specific malocclusion types. This enhances the system's applicability to a broader range of clinical cases.
In summary, the personalized 3D facial soft tissue landmark annotation system proposed in this study demonstrates significant efficiency, precision, and flexibility, making it a valuable tool in orthodontic clinical practice. Its high-precision annotations provide reliable support for preoperative diagnosis, postoperative comparative analysis, and precise localization of treatment targets, optimizing the treatment planning process. The flexibility of personalized templates enables the system to adapt to various complex cases, such as facial deformities or unique anatomical features, offering accurate annotation solutions for such cases. Additionally, the system's capability for batch processing 3D facial images supports large-scale studies in orthodontics, such as population feature analysis or facial morphology assessment, and its applications may extend to other related fields.
Limitation
Despite these promising results, the study has certain limitations. Firstly, the study population consisted exclusively of Chinese orthodontic patients from a single institution with a relatively young mean age, which introduces potential ethnic and age-related biases. Secondly, Craniofacial morphology exhibits significant ethnic variations, and the Basel Face Model (originally developed from predominantly Caucasian populations) may introduce systematic bias when applied to non-Caucasian ethnicities. Although our results demonstrated acceptable accuracy in the Chinese cohort, independent validation in diverse ethnic populations is necessary to confirm cross-ethnic generalizability.
Future research should focus on multi-center validation studies incorporating diverse ethnic populations and expanded age ranges, larger sample sizes for rare deformity types, and development of ethnicity-specific and age-specific template libraries to establish broader clinical applicability.
Conclusion
This study proposed a personalized automated 3D facial soft tissue landmarks annotation based on deep learning and computer vision. The method achieved efficient and precise annotation through standardized facial template construction, 3D facial nonlinear registration, and personalized landmark transfer. This system met clinical needs in orthodontics, providing significant convenience for the measurement of three-dimensional landmarks of soft tissues, the analysis of facial soft tissue morphology, and the diagnosis, treatment, and efficacy evaluation of malocclusion. Furthermore, the system was capable of constructing various specialized face templates, thereby enhancing the accuracy of annotating faces with severe deformities and broadening its clinical applicability. However, future research will focus on further optimizing the precise identification of landmarks in the inner and outer canthus regions and verifying their efficacy across diverse populations.
Data availability
The datasets generated and/or analysed during the current study are not publicly available due privacy or ethical restrictions, but are available from the corresponding author on reasonable request. The complete source code and configuration files used in this study will be made publicly available upon manuscript acceptance.
References
Luyten J, Vierendeel M, De Roo NMC, et al. A non-cephalometric three-dimensional appraisal of soft tissue changes by functional appliances in orthodontics: a systematic review and meta-analysis. Eur J Orthod. 2022;44(4):458–67.
D’Ettorre G, Farronato M, Candida E, et al. A comparison between stereophotogrammetry and smartphone structured light technology for three-dimensional face scanning. Angle Orthod. 2022;92(3):358–63.
Farkas LG, Bryson W, Klotz J. Is photogrammetry of the face reliable? Plast Reconstr Surg. 1980;66(3):346–55.
Ghoddousi H, Edler R, Haers P, et al. Comparison of three methods of facial measurement. Int J Oral Maxillofac Surg. 2007;36(3):250–8.
Weissler JM, Stern CS, Schreiber JE, et al. The Evolution of photography and three-dimensional imaging in plastic surgery. Plast Reconstr Surg. 2017;139(3):761–9.
Lübbers HT, Medinger L, Kruse A, et al. Precision and accuracy of the 3dMD photogrammetric system in craniomaxillofacial application. J Craniofac Surg. 2010;21(3):763–7.
Ten Harkel TC, Speksnijder CM, van der Heijden F, et al. Depth accuracy of the RealSense F200: low-cost 4D facial imaging. Sci Rep. 2017;7(1):16263.
Liu J, Zhang C, Cai R, et al. Accuracy of 3-dimensional stereophotogrammetry: comparison of the 3dMD and Bellus3D facial scanning systems with one another and with direct anthropometry. Am J Orthod Dentofacial Orthop. 2021;160(6):862–71.
Liu ZY, Yu J, Dai FF, et al. Three-dimensional changes in lip vermilion morphology of adult female patients after extraction and non-extraction orthodontic treatment. Korean J Orthod. 2019;49(4):222–34.
Lambros V, Amos G. Three-dimensional facial averaging: a tool for understanding facial aging. Plast Reconstr Surg. 2016;138(6):980e–2e.
Du F, Chen Q, Wang X, et al. Long-term facial changes and clinical correlations in patients with treated acromegaly: a cohort study. Eur J Endocrinol. 2021;184(2):231–41.
Serafin M, Baldini B, Cabitza F, et al. Accuracy of automated 3D cephalometric landmarks by deep learning algorithms: systematic review and meta-analysis. Radiol Med. 2023;128(5):544–55.
Zemouri R, Zerhouni N, Racoceanu D. Deep learning in the biomedical applications: recent and future status. Appl Sci. 2019;9(8):1526.
Berends B, Bielevelt F, Schreurs R, et al. Fully automated landmarking and facial segmentation on 3D photographs. Sci Rep. 2024;14(1):6463.
Yang Y, Zhang M, An Y, et al. Automated 3D perioral landmark detection using high-resolution network: artificial intelligence-based anthropometric analysis. Aesthet Surg J. 2024;44(8):Np606-np612.
Nishimura M, Tachiki C, Morikawa T, et al. Cranial vault deformation and its association with mandibular deviation in patients with facial asymmetry: A CT-based study. Diagnostics (Basel). 2025;15(13):1702.
Bollhalder J, Hänggi MP, Schätzle M, et al. Dentofacial and upper airway characteristics of mild and severe class II division 1 subjects. Eur J Orthod. 2013;35(4):447–53.
Kerr WJ, Miller S, Dawber JE. Class III malocclusion: surgery or orthodontics? Br J Orthod. 1992;19(1):21–4.
A 3D Face Model for Pose and Illumination Invariant Face Recognition. 2009 Sixth IEEE International Conference on Advanced Video and Signal Based Surveillance; 2009: 296–301.
Jabberi M, Wali A, Chaudhuri BB, Alimi AM. 68 landmarks are efficient for 3D face alignment: what about more?: 3D face alignment method applied to face recognition. Multimed Tools Appl, 2023: 1-35.
Wu Y, Ji Q. Facial landmark detection: a literature survey. Int J Comput Vis. 2019;127(2):115–42.
Iterative Reweighted Local Cross Correlation Method for Nonlinear Registration of Multiphase Liver CT Images. 2021 IEEE International Conference on Image Processing (ICIP); 2021: 136–140.
Loshchilov I, Hutter F: Decoupled Weight Decay Regularization. In: International Conference on Learning Representations: 2017; 2017.
Liu ZY, Chen G, Dai FF, et al. Analysis of correlation of 3-dimensional lip vermilion morphology and dentoskeletal forms in young Chinese adults on the basis of sex and skeletal patterns. Am J Orthod Dentofacial Orthop. 2021;159(5):e423–37.
Burger J, Blandano G, Facchi GM, Lanzarotti R. 2S-SGCN: a two-stage stratified graph convolutional network model for facial landmark detection on 3D data. Comput Vis Image Underst. 2025;250:104227.
Torosdagli N, Liberton DK, Verma P, et al. Deep geodesic learning for segmentation and anatomical landmarking[J]. IEEE Trans Med Imaging. 2019;38(4):919–31.
Bermejo E, Taniguchi K, Ogawa Y, et al. Automatic landmark annotation in 3D surface scans of skulls: Methodological proposal and reliability study. Comput Methods Programs Biomed. 2021;210:106380.
Jawahar CV, Li H, Mori G,Schindler K editors. Multi-view Consensus CNN for 3D Facial Landmark Placement. In: Computer Vision – ACCV 2018: 2019// 2019; Cham: Springer International Publishing; 2019: 706–719.
Fagertun J, Harder S, Rosengren A, et al. 3D facial landmarks: inter-operator variability of manual annotation. BMC Med Imaging. 2014;14:35.
Hajeer MY, Ayoub AF, Millett DT, et al. Three-dimensional imaging in orthognathic surgery: the clinical application of a new method. Int J Adult Orthodon Orthognath Surg. 2002;17(4):318–30.
Al-Baker B, Alkalaly A, Ayoub A, et al. Accuracy and reliability of automated three-dimensional facial landmarking in medical and biological studies. A systematic review. Eur J Orthod. 2023;45(4):382–95.
Galvánek M, Furmanová K, Chalás I, Sochor J. Automated facial landmark detection, comparison and visualization. In. Automated facial landmark detection, comparison and visualization. Proceedings of the 31st Spring Conference on Computer Graphics; Smolenice, Slovakia: Association for Computing Machinery; 2015.
Chong Y, Du F, Ma X, et al. Automated anatomical landmark detection on 3D facial images using U-NET-based deep learning algorithm. Quant Imaging Med Surg. 2024;14(3):2466–74.
Zhang Y, Xu Y, Zhao J, et al. An automated method of 3D facial soft tissue landmark prediction based on object detection and deep learning. Diagnostics (Basel). 2023;13(11):1853.
Baksi S, Freezer S, Matsumoto T, Dreyer C. Accuracy of an automated method of 3D soft tissue landmark detection. Eur J Orthod. 2021;43(6):622–30.
Wang K, Zhao X, Gao W, Zou J. A coarse-to-fine approach for 3D facial landmarking by using deep feature fusion. Journal. 2018;10(8):308.
Acknowledgements
Not applicable.
Funding
This study was funded by the National Natural Science Foundation of China No.82370999. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Author information
Authors and Affiliations
Contributions
W.Y.W. contributed to design, data acquisition, analysis, and interpretation, drafted and critically revised the manuscript. Z.Z.C.contributed to design, data acquisition, analysis, and interpretation, drafted and critically revised the manuscript. L.J.L.contributed to data acquisition and interpretation, performed statistical analyses, critically revised the manuscript. Y.Q.H. contributed to data acquisition and interpretation, critically revised the manuscript. H.X.Y. contributed to data acquisition and interpretation, critically revised the manuscript. Z.X.contributed to data acquisition and interpretation, critically revised the manuscript. S.C.Y. contributed to data acquisition and interpretation. Y.J.contributed to conception, carried out the computer part of the study, critically revised the manuscript. G.J. contributed to conception, interpretation of data, critically revised the manuscript. All authors read and approved the final manuscript.
Corresponding authors
Ethics declarations
Ethics approval and consent to participate
All participants provided written informed consent prior to enrollment in the study. This research received approval from the Ethics Committee of Shandong University (approval number: 20241006) and adheres strictly to the ethical principles for human research outlined in the Declaration of Helsinki (2013 revised edition).
Consent for publication
All participants provided written informed consent for the publication of their anonymized clinical data and any potentially identifiable images.
Competing interests
The authors declare no competing interests.
Additional information
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary Information
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
About this article
Cite this article
Wu, Y., Zhao, Z., Lu, J. et al. A personalized automated system of 3D facial soft tissue landmarks annotation based on deep learning and computer vision. BMC Oral Health 26, 6 (2026). https://doi.org/10.1186/s12903-025-07448-3
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1186/s12903-025-07448-3







