Picking Winners in Dhaka Stock Exchange Without Getting Tailed
A forecasting contest cannot identify the best tail-risk model if the specification favoured by the descriptive-fit comparison is never allowed to compete. This study examines a selection–evaluation gap and shows that including the omitted specification changes the preferred model family in the present application. Using 2,383 synchronised daily returns on the DSEX, DS30, and DSES indices of the Dhaka Stock Exchange from 3 January 2016 to 19 July 2026, this study runs a single, fully specified out-of-sample competition among GARCH-t, GJR-t, GJR-EVT, EGARCH-t, and EGARCH-EVT across 744 one-day-ahead forecasts during 2023–2026. Every competitor — econometric and algorithmic alike — is estimated on an expanding window, refitted on the same four annual dates, held fixed within each forecast year, and evaluated using information available only through t-1. EGARCH-t achieves the “best” descriptive BIC for all three indices, and most importantly, in all 24 matched comparisons, the EGARCH family also achieves lower values of the pinball loss than the GJR alternatives. In the nine comparisons where Diebold–Mariano tests reject equal predictive accuracy, they never favour GJR. The ranking is not purchased through undercoverage since all five econometric specifications pass the unconditional-coverage and independence tests and EGARCH-EVT achieves one per cent VaR which is 2.44 to 6.43 basis points less conservative in calm states and 8.88 to 14.33 basis points less conservative in turbulent states with identical numbers of violations in each cell of the index and state matrix. Nevertheless, a histogram gradient-boosting quantile learner is fitted with fixed settings and lagged market features and refitted on the same real-time schedule and fails unconditional coverage on both of the tail levels for each index showing a calibration failure for this specification. If the contest was restricted to GARCH and GJR then the wrong family was crowned in every cell. Because the empirical results are limited to one exchange, with three indices that have overlapping constituents and share the same institutional environment, the design principle that they derive is general.