This project studies whether stochastic implementation choices in modern machine learning materially affect empirical asset pricing conclusions. The first paper, Random Seeds and Out-of-Sample Performance in Empirical Asset Pricing via Machine Learning, documents that models such as Random Forests, XGBoost, and feed-forward neural networks can produce different forecasts, portfolio holdings, Sharpe ratios, CAPM alphas, and out-of-sample R2 values when trained on the same data, using the same sample split, model class, hyperparameter protocol, and portfolio construction, but with different random seeds. The current evidence is based on monthly U.S. equities from 1972–2022, more than 150 firm-level characteristics, and 500 stochastic runs/ensemble draws across Random Forest, XGBoost, and neural networks with 1 to 5 hidden layers. The results show that the random seed is not merely a technical detail: algorithmic randomness can propagate through tuning, forecasts, rankings, portfolio composition, feature importance, and ultimately economic performance.
The requested extension will allow the project to move from documenting conditional seed-induced uncertainty in the 2018–2022 out-of-sample period to assessing whether this uncertainty is persistent across market regimes. A central next step is to extend the out-of-sample evaluation to earlier and longer test periods, including rolling OOS windows and block-bootstrap analyses. This will separate variation due to random seed choice from variation due to the particular realized test sample, and will show whether seed sensitivity is a general property of machine-learning asset pricing or specific to the recent 2018–2022 period.
The project is also developing into a broader research agenda on reproducibility in financial machine learning. A second paper will replicate influential machine-learning asset-pricing studies and rerun their stochastic models across many random seeds under otherwise fixed protocols. The purpose is to quantify how much published conclusions about predictive accuracy, Sharpe ratios, alphas, feature importance, and model rankings depend on unreported or underreported seed choices. This will provide evidence on the potential scope for opportunistic seed selection, or “seed-hacking,” in academic finance, while also proposing transparent reporting standards for stochastic machine-learning studies.
The project contributes to empirical asset pricing, financial machine learning, and research transparency. Its outputs will include distributional performance evidence, diagnostics for forecast and portfolio stability, guidance on minimum seed reporting, and replication-based evidence on the robustness of major results in the literature. High-performance computing resources are essential because each extension requires repeated hyperparameter tuning, rolling refits, and portfolio evaluation across many stochastic model realizations and long historical samples.
Main supervisor: Gustav Martinsson, Professor of Finance, Stockholm University.