Flow matching is used in a growing array of scientific domains (e.g. protein and molecule design and optimization), and starting to show promise in large language modeling. Current flow matching frameworks require the length of the sequence (protein, molecule, sentence, etc) to be specified in advance. Recently, we developed a new flow matching framework, called "Branching Flows" (https://arxiv.org/abs/2511.09465 and see https://murrellgroup.github.io/BranchingFlows/ for a project page), that allows the number of elements in the flow state to change during design. Briefly, in Branching Flows the elements in the state evolve over a forest of binary trees, branching and dying stochastically. These rates are learned by the model, allowing it to control, during generation, the number of elements in the sequence. Variable length generation solves standing problems with standard models, especially in conditional generation. For example, if you wish to design a molecule that binds to a pocket you typically do not know the number of atoms required in advance, and Branching Flows models solve this directly. This works whether the state is continuous, discrete, or manifold-valued, making this an attractive framework for generative modeling in scientific domains.
With the mathematics of this established, and with clear demonstration of promise in a number of small-scale modeling tests across molecule generation, protein generation, and amino acid sequence models, we are now looking to expand and scale a small number of promising protein design use cases, and pilot finetuning of Branching Flows LLMs (up to 1B parameters).
This project will:
1) Run short pilot finetuning runs on three larger protein design models, selecting one to move forward for a full-scale training run.
2) Train, from scratch, an atom-level model for joint protein and molecule confirmation sampling and design.
3) Explore short Branching Flows finetuning of a number of <1B parameter text LLMs, including autoregressive and diffusion/flow matching LLMs as a starting point.
The compute will also be used for inference benchmarks and comparisons to other approaches required for publication.