Shadow-based geolocation from synthetic imagery
Predicting GPS coordinates and UTC time from 512×512 rendered scenes, scored on a composite of Haversine distance and circular time error. Analytical solar-position solvers were prohibited, so the model had to learn the geometry rather than be told it.
The counter-intuitive result was that bigger backbones lost. ResNet-101 and a fine-tuned ViT both underperformed ResNet-34 and ResNet-50 on this task. The final submission was a three-seed ensemble over those two architectures, with a different validation split per seed, averaging predictions in xyz space for position and sin/cos space for time rather than averaging raw angles.