🦷 3D Teeth Landmark Detection

A Point-Transformer-v3-style network that detects six classes of anatomical dental landmarks directly on 3D intraoral scan point clouds — no tooth segmentation stage. Built during a research internship at the National Dental Centre of Singapore and evaluated under the official 3DTeethLand (MICCAI 2024) protocol.

Model output on a held-in example scan

Arch
Show

Drag to rotate · scroll to zoom · click legend entries to isolate a landmark class. Enable “Ground truth” (hollow diamonds) to compare against the clinician annotations.

Accuracy on this scan

ClassPredGTMedian err< 2 mm

Error = distance from each predicted landmark to the nearest ground-truth landmark of the same class.

Official challenge results

ArchmAPmAR
Lower0.5360.412
Upper0.5530.438

Scored with the challenge's own evaluation code on the 50-scan held-out test split per arch.

Tolerance0.5 mm1.0 mm1.5 mm2.0 mm
mean AP0.0910.4730.7010.790
mean AR0.2630.6350.7870.843

Lower arch, test split.

How it works

  1. Input — 10,000 vertices sampled from the scan mesh, each with its surface normal: [N, 6] features (x, y, z, nx, ny, nz).
  2. Encoder — 4 stages × 2 pre-norm transformer blocks (multi-head self-attention + MLP) at hidden dim 256. Unlike standard PTv3, no downsampling is applied, so per-point resolution survives end to end.
  3. Distance-map regression — a per-point MLP predicts, for each of the six classes, the normalized distance to the nearest landmark of that class (clamped at 15 mm, sqrt-sharpened). Dense targets give every point a gradient — this trained far better than the sparse Gaussian heatmaps tried first.
  4. Post-processing — points below a per-class threshold are clustered with DBSCAN; each cluster yields one landmark. Handles variable landmark counts (missing teeth) without knowing the number in advance.