NAISS
SUPR
NAISS Projects
SUPR
Pinball: Scaling A Hierarchical Graph Transformer for Efficient Long-Context Genomics Modeling
Dnr:

NAISS 2026/3-707

Type:

NAISS Medium

Principal Investigator:

Martin Enge

Affiliation:

Karolinska Institutet

Start Date:

2026-08-29

End Date:

2027-03-01

Primary Classification:

10610: Bioinformatics and Computational Biology (Methods development to be 10203)

Secondary Classification:

10210: Artificial Intelligence

Tertiary Classification:

10609: Genetics and Genomics (Medical aspects at 30107 and agricultural at 40402)

Allocation

Abstract

Transformers have achieved strong performance across sequence modeling tasks, but their application to long-context modeling remains constrained by the quadratic complexity of global attention. We introduce Pinball, a hierarchical graph-based Transformer architecture that combines local token-level attention with multi-scale message passing to enable efficient long-range communication. By decoupling local computation from global information flow, Pinball scales approximately linearly with sequence length while preserving full-resolution token modeling. Initial results show competitive performance on language modeling benchmarks, stable behavior at extended context lengths, strong performance on genomic sequence prediction tasks, and improved copy fidelity across distances beyond the local attention window. Together, these results suggest that hierarchical interaction structures provide a practical and generalizable alternative to full attention for long-context sequence modeling. The potential impact is particularly strong in genomics, where many regulatory effects depend on nucleotide-resolution sequence context spanning tens to hundreds of kilobases or more. Current state-of-the-art genomic models remain computationally expensive to train and deploy at such resolutions. Pinball’s approximately linear scaling offers a path toward efficient, full-resolution genomic sequence models that can be trained and used by a broader research community, supporting applications in regulatory genomics, variant effect prediction, enhancer–promoter interaction modeling, and long-range sequence design. The architecture is now mature and has been validated in small-scale local experiments. However, demonstrating its full potential requires cloud compute to benchmark Pinball at scale across model sizes, sequence lengths, and training modalities. Access to cloud resources would allow us to systematically evaluate scaling behavior, compare against established Transformer and genomics baselines, train larger and more general models, and extend the architecture toward high-impact biomedical applications, including cancer genomics, and spatial genomics. In particular, we aim to model long-distance regulatory effects, predict the functional consequences of noncoding variants, identify enhancer–promoter dependencies, and explore large-context sequence design beyond current state-of-the-art limits. These capabilities have been demonstrated in preliminary form, but further progress is currently bottlenecked by compute rather than architecture or implementation readiness. Cloud compute would therefore directly convert an already functional prototype into a scalable, benchmarked, and broadly useful platform for long-context sequence modeling in genomics and beyond.