Hagfish represent one of the two surviving lineages of jawless vertebrates and occupy a key phylogenetic position for reconstructing the early evolution of vertebrate genomes and immune systems. Together with lampreys, hagfish possess an alternative adaptive immune system based on variable lymphocyte receptors (VLRs), rather than the immunoglobulin/T-cell receptor system of jawed vertebrates. However, genomic resources for hagfish remain limited, restricting detailed analysis of their immune gene repertoire, genome organization, and cell-type-specific transcriptional programs.
This project aims to generate a high-quality, haplotype-resolved somatic genome assembly and comprehensive gene annotation for the Atlantic hagfish, Myxine glutinosa, using material collected through our collaboration with the Michael Sars Centre and the University of Bergen Marine Biological Station. PacBio HiFi sequencing data from a single individual will be assembled using hifiasm, followed by haplotype resolution, duplicate removal, contamination screening, and chromosome-scale scaffolding using publicly available Hi-C data. As the Hi-C data originate from a different individual, they will primarily be used for scaffolding and structural validation rather than chromosome-level phasing.
The hagfish genome presents a substantial computational challenge because of its large size (approximately 3.5–4 Gb) and exceptionally high repeat content. Preliminary repeat modelling indicates that approximately 80% of the genome may consist of repetitive sequences, with a large fraction remaining unclassified. Consequently, iterative repeat discovery, classification, masking, assembly evaluation, and manual curation will require considerable CPU time, memory, and storage.
Gene annotation will integrate bulk RNA-seq and de novo transcriptome assemblies from multiple tissues, together with protein evidence, using workflows including BRAKER3. Both haplotypes will be annotated and functionally characterized. Particular attention will be given to genes involved in VLR-based adaptive immunity and hematopoiesis, for which conventional automated annotation can be problematic because of unusual gene structures and short coding regions. The resulting genome will also provide the reference required for interpretation of our newly generated single-cell transcriptomic data from hagfish peripheral blood.
NAISS computational resources will therefore enable the assembly, repeat analysis, annotation, and integration of large genomic and transcriptomic datasets required to establish a robust genomic reference for this evolutionarily important vertebrate.