Turning raw data into published genomes faster with the Australian Tree of Life Genome Engine

The Black-Eared Miner genome has become openly available as part of the effort to uplift the generation, publication and re-use of reference genomes for Australian species. Like a lot of sequence data, the bird’s genome remained unassembled until now because the process of transforming raw sequence data into genome assemblies for downstream analyses is time-intensive and computationally expensive. The Australian Tree of Life (AToL) Genome Engine changes the equation by enabling rapid, automated genome assembly, annotation and publication.

Sequenced as part of the Threatened Species Initiative in 2023, the Black-Eared Miner genome was recently ingested into the AToL Genome Engine for assembly and rapidly deposited in the European Nucleotide Archive (ENA) to achieve maximum reach  and utility of the data. 

Under development from the AToL Bioinformatics team at BioCommons, the Genome Engine accelerates the assembly and annotation of genomic data generated by Bioplatforms Australia’s Framework Initiatives. The first of the Australian Venom Innovation and Discovery Initiative’s (AVID) genome assemblies, describing the native bag-shelter moth or processionary caterpillar, Ochrogaster lunifer, has also been submitted to ENA. AVID data is currently only available to researchers in the AVID Initiative, but will become publicly available after an embargo period. The code base used to generate each assembly is also published in GitHub.

How does the Genome Engine work?

The Genome Engine is a semi-automated workflow for assembling and annotating genome sequences from raw sequence data, brokering data to International Nucleotide Sequence Database Collaboration (INSDC) repositories.

This involves:

  • Ingesting raw sequence data from the Bioplatforms Australia Data Portal

  • Processing sampling and sequencing metadata

  • Assembling genome sequences from sequence read data

  • Annotating assembled genomes

  • Brokering sample metadata, sequence reads, and genome assemblies to the ENA

  • Providing details and metrics about sampling, sequencing and assembly

Infographic illustrating the automation steps for the AToL Genome Engine

How the AToL Genome Engine works

Who can use the Genome Engine?

So far, the Genome Engine has been refined by exposing it to the needs of the Bioplatforms Framework Initiatives. These large research consortia have provided diverse genomes representing Australian fish and birds, threatened species, and venomous animals. While it is configured to ingest data available from the Bioplatforms Australia Data Portal, the Genome Engine will eventually be made available to all Australian researchers.

If you have interesting sequencing data of a species significant to Australia, get in touch to see if you can benefit from this expanding scope. Otherwise stay tuned for when the AToL Genome Engine launches as a service open to all Australian researchers.


View the Black-Eared Miner genome draft assembly 

Find out more about the AToL Genome Engine 

Contact the team if you have a genome sequence of an Australian species needing assembly: atol-bioinformatics@unimelb.edu.au


AToL Bioinformatics is co-funded by Bioplatforms Australia and Minderoo Foundation, with project partners at the University of Melbourne, QCIF, the European Nucleotide Archive, and the Darwin Tree of Life.

Next
Next

Strengthening global biodata sustainability: Australia joins international advisory leadership