Many scientific optimization problems are characterized by very large design spaces combined with expensive or otherwise limited evaluations. Machine-learning models can substantially reduce this cost by learning predictive representations from previously evaluated candidates and using them to guide subsequent search. However, as optimization progresses, the evaluated data increasingly differ from the distributions on which the models were originally trained, which can reduce predictive reliability and consequently the efficiency of the search.
In this project, I will develop machine-learning methods for data-efficient iterative optimization under such distribution shifts. The work will combine generative modelling, predictive neural networks, uncertainty estimation and adaptive data selection in closed learning loops. The methods will be evaluated on established scientific optimization benchmarks, with particular emphasis on how quickly useful solutions can be identified as a function of the available evaluation budget.
The computational work will involve repeated training and evaluation of neural-network models across different tasks, sampling strategies and random seeds, followed by comparisons with established active-learning and optimization methods.
The overall aim is to improve how limited computational or experimental evaluations are allocated in large scientific design spaces, enabling more efficient search when accurate labels are expensive to obtain.