Table 4 Ablation study of each module
From: Transferable adversarial masked self-distillation for unsupervised domain adaptation
\({{\mathcal {L}}_{\textrm{cls}}}({x^s},{y^s})\) | \({{\mathcal {L}}_{\textrm{adv}}}({x^s},{x^t})\) | \({{\mathcal {L}}_{{\text {self-KD}}}}({x^t};\theta )\) | \({{\mathcal {L}}_{{\text {MIM}}}}({x^t};\theta )\) | \({{\mathcal {L}}_{\textrm{patch}}}({x^s},{x^t})\) | \(I({p^t};{x^t})\) | A \(\rightarrow \) W | D \(\rightarrow \) W | W \(\rightarrow \) D | A \(\rightarrow \) D | D \(\rightarrow \) A | W \(\rightarrow \) A | Avg |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
\(\checkmark \) | 89.18 | 98.87 | 100 | 88.76 | 80.09 | 79.77 | 89.45 | |||||
\(\checkmark \) | \(\checkmark \) | 90.12 | 98.89 | 100 | 90.45 | 83.43 | 83.23 | 91.02 | ||||
\(\checkmark \) | \(\checkmark \) | \(\checkmark \) | 93.14 | 98.93 | 100 | 93.32 | 84.63 | 85.13 | 92.53 | |||
\(\checkmark \) | \(\checkmark \) | \(\checkmark \) | \(\checkmark \) | 95.34 | 99.03 | 100 | 94.14 | 84.93 | 85.67 | 93.19 | ||
\(\checkmark \) | \(\checkmark \) | \(\checkmark \) | \(\checkmark \) | \(\checkmark \) | 95.44 | 99.01 | 100 | 94.69 | 85.05 | 85.96 | 93.36 | |
\(\checkmark \) | \(\checkmark \) | \(\checkmark \) | \(\checkmark \) | \(\checkmark \) | \(\checkmark \) | 96.48 | 99.5 | 100 | 96.99 | 85.87 | 86.26 | 94.18 |
\({{\mathcal {L}}_{\textrm{TAMS}}}\) with traditional self-attention in ViT encoder | 93.52 | 98.79 | 100 | 94.19 | 84.33 | 85.79 | 92.77 | |||||
\({{\mathcal {L}}_{\textrm{TAMS}}}\) with the cross-attention mechanism in ViT encoder | 95.49 | 99.13 | 100 | 95.82 | 85.28 | 85.98 | 93.62 | |||||
\({{\mathcal {L}}_{\textrm{TAMS}}}\) with the weighted cross-attention mechanism (Eq. (6)) in ViT encoder | 96.48 | 99.5 | 100 | 96.99 | 85.87 | 86.26 | 94.18 | |||||