Previously, we developed and implemented transformer models that use Matryoshka Representation Learning (MRL), a technique that allows for optimisation of multiple sizes of the same model simultaneously. MRL produces multiple usable model sizes which can be selected at inference. Our results showed that MRL improves the compute-performance trade-off of transformers by (i) improving efficiency exponentially as a function of model size choice while (ii) causing only minor increases in loss.
In our first small compute project, we successfully established the foundation for a research paper to be submitted to an international and competitive research forum. In order to complete our tests, we would like to hereby apply for a shorter small compute project, which will then let us finish our manuscript and get it ready for submission.
This new project aims to assess the performance of LLMs trained on an MRL objective.