The impact of socio-economic data on deep learning crop yield prediction in Cameroon

Most crop yield models in the literature stack whatever data is available and report a single score. This study asks a narrower and more useful question: what does each data modality actually contribute?

The design uses three nested datasets — satellite indices alone (D1), satellite plus climate (D2), and satellite plus climate plus socio-economic indicators (D3) — held constant across six deep learning architectures: CNN-LSTM, ConvLSTM, ConvLSTM-ViT, CNN-GRU, TCN and CNN-Transformer. Crops are maize, cassava and cocoa; regions are Centre and Nord, Cameroon.

Because the datasets are nested rather than parallel, the difference between D2 and D3 is attributable to the socio-economic modality rather than to a change of model or of split. The pipeline enforces target scaling applied after the split, scalers fitted on training folds only, and early stopping on a held-out set.

Status: manuscript in preparation. Target venues: Computers and Electronics in Agriculture; Remote Sensing (MDPI).