As discussed, vast numbers of candidate datasets exist that could be related to the problem of COVID19 forecasting. However, these datasets cannot be used indiscriminately. We select data sources based on whether they could have a predictive signal for the disease outcomes. Selecting multiple datasets from the same class of causes can obfuscate their predictive power. Therefore, we select datasets, one each from the classes of econometrics, demographics, mobility, non-pharmaceutical interventions, hospital resource availability, historical air quality. From each of these datasets, we further select covariates that could have an impact on the model compartments. We allow covariates to influence only those compartments (and hence transition rates) on which we posit that there exists a causal relationship (Table 1).