This project aims to characterize WOX4 transcript diversity and regulatory conservation using public long-read transcriptome datasets and comparative plant genomics. The original work focused on identifying canonical and alternative WOX4 transcripts, including condition-associated transcription start sites, across long-read RNA-sequencing datasets from different plant materials and treatments.
The project has subsequently expanded to a targeted comparative analysis across approximately 284 plant genomes. This analysis evaluates the conservation of the WOX4 gene body, the first coding intron, and adjacent regulatory regions. The workflow includes retrieval and normalization of genome assemblies and annotations, ortholog identification, protein and nucleotide similarity searches, gene-structure validation, sequence extraction, multiple-sequence alignment, and manual quality control.
The results will provide an evolutionary framework for identifying conserved candidate regulatory elements and will guide downstream molecular validation in Arabidopsis and other plant species.