Follow a data engineer’s journey in building the infrastructure needed to support a data sharing platform. We’ll cover key concepts in managing the end to end flow of data, and hear what considerations and choices are essential to achieving a robust system.
The fundamental questions that must be answered before any new platform is built will be explored, along with the design elements that significantly impact the effective use of data platforms like a data commons. This presentation will use the open-source Gen3 data platform as an example, drawing on the real-life experience of creating the Australian Cardiovascular disease Data Commons.
This presentation will step you through important elements of data operations including:
The many functions of the foundational data model
How data model changes impact the way data flows through the entire system
The role of the data pipeline in managing provenance, validation, data quality, platform governance, and reproducibility
How tooling maps onto each stage
The webinar will be delivered in two distinct halves, and you can choose to leave after the presentation on concepts, or stay online for the second hour about the practical application of this framework.
The second part of this webinar will provide a live end to end demonstration of how to build and operate a data pipeline for Gen3 in AWS, using a test data commons that Australian BioCommons has been building with health, research and data infrastructure collaborators. The run-through will dive deep into: building a data model, configuring and deploying infrastructure, setting up the CLI tooling, deploying the model, generating synthetic data, uploading and deleting metadata, cutting data releases, managing transformations, and inspecting data operations in the system.
Speaker: Dr Joshua Harris, Data Engineer, Human Genome Informatics, Australian BioCommons
Date/Time: 02 September 2026, 1 - 3 pm AEST / 12:30 - 2:30 pm ACST / 11 am - 1 pm AWST (check in your timezone)
Who the webinar is for:
This webinar is for anyone that is interested in building data sharing platforms for research, and wants to understand how to build robust, production grade systems that support data provenance, governance, and reproducibility.
How to join:
This webinar is free to join but you must register for a place in advance.
--------
This event is part of a series of bioinformatics training events. If you’d like to hear when registrations open for other events, please subscribe to the Australian BioCommons newsletter.