£400Registration Fee
Register Now- Overview
- Instructors
- Schedule
Course Description
This course will cover the metagenomics data analysis workflow from data generation to downstream analysis. Participants will explore the tools to manage, share, analyze and interpret metagenomics data. The content will include issues of data quality control and how to process the data until publication.
Both targeted marker-gene (e.g., 16S) and whole-genome shotgun (WGS) approaches will be covered. Discussions will also explore considerations when selecting the sequencing options (long or short reads), exploring targeted data to OTUs or ASVs, assembling and binning metagenomics data, available tools, and the downstream analyses that can be carried out.
What You’ll Learn
- Define the core concepts, terminology, and scope of metagenomics, and distinguish it from related fields such as genomics, culturomics, and single-cell sequencing.
- Explain the biological and ecological rationale for studying microbial communities in situ, including the concept of the “unculturable majority.”
- Describe key considerations in experimental design for metagenomic studies, including replication, controls (negative/positive), and confounding variables.
- Recognize appropriate sample collection, storage, and DNA/RNA extraction methods for different environments (e.g., soil, gut, water) and their impact on downstream results.
- Interpret and compare the two main methodological approaches (targeted vs shotgun) metagenomics including their respective advantages, limitations, costs, and appropriate use cases.
- Identify relevant sequencing platforms (e.g., Illumina, Oxford Nanopore, PacBio) and how platform choice influences experimental design and analysis strategy.
- Execute quality control and pre-processing of raw sequencing data, including adapter trimming, quality filtering, host/contaminant removal, and assessment of sequencing depth and coverage.
- Explain the principles behind key computational steps such as read assembly, binning, taxonomic classification, and functional annotation.
- Evaluate the outputs of metagenomic pipelines critically, including assessment of assembly quality, bin completeness/contamination and classification confidence.
- Describe downstream analysis options, including diversity analysis (alpha/beta diversity), differential abundance testing, functional pathway analysis, and community comparison across conditions.
- Apply appropriate statistical methods and visualization techniques to interpret and communicate metagenomic results.
- Apply FAIR data principles and other good practices in bioinformatics to metagenomic datasets, including proper deposition in public repositories.
.
Course Format
Flexible Learning Structure
Learn through a carefully structured mix of lecture recordings and guided exercises that you can pause, revisit, and complete at your own pace—ideal for busy professionals or those balancing multiple commitments.
Access Anytime, Anywhere
All course content is available on-demand, making it accessible across all time zones without the need to attend live sessions or adjust your schedule.
Independent Exploration with Support
Engage deeply with course topics through self-directed study, with the option to reach out to instructors via email for clarification or deeper discussion.
Comprehensive Learning Resources
Gain full access to the same high-quality materials provided in live sessions, including code, datasets, and presentation slides—all available to download and keep. Please note recordings can only be streamed.
Work With Your Own Data, On Your Terms
Apply what you learn directly to your own data projects as you go, allowing for a personalized and immediately practical learning experience.
Continued Guidance and Resource Access
Receive 30 days of post-enrolment email support and unrestricted access to all session recordings during that time, so you can review and reinforce your learning as needed.
Who Should Attend / Intended Audiences
This course is designed for PhD students, postdoctoral researchers, and other life scientists from academia and industry who are working in metagenomics and are in the early stages of data preparation or analysis.
To get the most from the course, participants should have knowledge equivalent to the learning outcomes of First Steps with R in Life Sciences and the UNIX Fundamentals e-learning course.
Specifically, participants should be familiar with Rstudio, how to install a library, matrix and data frame manipulation, import and export data from text files. In case of doubt, evaluate your R skills with this quiz, before registering.
Equipment and Software requirements
You will be required to create a GitHub account at: https://github.com/signup. During the course, we will use the free tier of GitHub Codespaces in order to count with a standardized working environment for all the participants.
⚠ WARNING! Do not start using the Codespaces until the day of the course since it has limited computing hours and storage space, and they can be run out pretty fast.
A working webcam is recommended to support interactive elements of the course. We encourage participants to keep their cameras on during live Zoom sessions to foster a more engaging and collaborative environment.
While not essential, using a large monitor—or ideally a dual-monitor setup—can significantly enhance your learning experience by allowing you to view course materials and work in R simultaneously.
All necessary packages will be introduced and installed during the workshop. A comprehensive list of required packages will also be shared with participants ahead of the course to allow for optional pre-installation.
Dr. Nikolay Oskolkov
Nikolay is a bioinformatician, computational biologist, and data scientist working at the intersection of biology, medicine, statistics, and artificial intelligence. His research focuses on applying mathematical statistics, machine learning, and deep learning methods to complex biological and biomedical datasets, including genomics, transcriptomics, microbiome research, single-cell data, metagenomics, and multi-omics integration.
Nikolay has a PhD in theoretical physics from 2007, he transition to the Life Sciences in 2011. He currently leads the Metabolic Research Group (MRG) within the TARGETWISE project at the National Institute of Research and Innovation in Latvia, and having a teaching position at Lund University, Sweden, he has previously held research positions at the Danish Technical University, University of North Carolina, Lund University and the National Bioinformatics Infrastructure Sweden (NBIS/SciLifeLab).
Nikolay has more than 20 years of teaching experience and is widely recognised for his ability to communicate advanced statistical and computational methods to researchers from diverse scientific backgrounds. His expertise spans both frequentist and Bayesian statistics, machine learning, dimensionality reduction, clustering, bioinformatics, and scientific programming in R and Python. He has delivered numerous international workshops, summer schools, and professional training courses in computational biology, genomics, and AI-driven biomedical research.
Education & Career
- PhD in Theoretical Physics (2007)
• Transitioned from theoretical physics to bioinformatics and computational biology in 2011
• Group Leader (PI), Metabolic Research Group, TARGETWISE Project, Latvia
• Former researcher and bioinformatician at Lund University and NBIS/SciLifeLab, Sweden
• Author of more than 60 peer-reviewed scientific publications with extensive international collaborations in computational biology and biomedical research
Research Focus
Nikolay’s work centres on extracting biological insight from large-scale, high-dimensional datasets using advanced statistical and machine learning approaches. His research interests include:
- Machine learning and deep learning for biomedical and omics data
• Multi-omics integration and systems biology
• Single-cell transcriptomics and dimensionality reduction methods
• Population genomics and evolutionary biology
• Microbiome, environmental DNA, and ancient DNA analysis
• Statistical modelling and Bayesian approaches for complex biological systems
• AI applications in precision medicine and drug discovery
Current Projects
- Development of machine learning methods for multi-omics data integration and drug discovery in metabolic diseases
• AI-driven approaches for genomics and computational biology
• Statistical and computational methods for ancient and environmental DNA research
• Machine learning analysis workflows for single-cell and population genomics datasets
• Research on metabolic diseases through integrative bioinformatics and systems biology approaches
Professional Consultancy
Nikolay provides expert consultancy in biological and biomedical data analysis, supporting academic researchers, healthcare scientists, and industry teams. His consultancy expertise includes:
- Bioinformatics and computational biology
• Medical genomics and precision medicine
• Single-cell and multi-omics data analysis
• Metagenomics and population genomics
• Frequentist and Bayesian statistical modelling
• Machine learning and deep learning applications
• Scientific programming in R, Python, Bash, and C++
• Study design, data analysis pipelines, and reproducible research workflows
Teaching & Skills
- More than 20 years of teaching experience in statistics, machine learning, and computational biology
• Teaches topics including machine learning, deep learning, Bayesian statistics, dimensionality reduction, clustering, single-cell analysis, genomics, and bioinformatics
• Instructor for international courses and workshops through organisations including Instats, Physalia, NBIS SciLifeLab, TARGETWISE, and RaukR
• Strong advocate for rigorous statistical thinking, reproducible research, and accessible scientific education
• Experienced in translating advanced computational methods into practical tools for life scientists and healthcare researchers
Links
Session 1 – 02:30:00 – Introduction to Metagenomics, goals and opportunities
Overview of metagenomics concepts and history; culture-independent approaches, study design considerations, and applications across environmental, clinical, and biotechnological fields; introduction to sequencing technologies, short and long reads.
Break – 00:30:00
Session 2 – 02:30:00 – Quality Control and Contamination Assessment
Quality assessment of sequencing data; adapter removal, filtering strategies, contamination detection, and best practices for genomic and transcriptomic analyses.
Session 3 – 02:30:00 – Targeted sequencing I (Introduction and theory)
Principles of amplicon-based sequencing; marker gene selection (16S, ITS), primer design, PCR amplification biases, and library preparation strategies for targeted metagenomics; OTUs vs ASVs.
Break – 00:30:00
Session 4 – 02:30:00 – Targeted sequencing II (Downstream analyses)
Processing of amplicon sequencing data; OTU/ASV inference, taxonomic assignment, diversity metrics, and statistical approaches for microbial community comparison.
Session 5 – 02:30:00 – Whole genome sequencing and taxonomic classification
Principles of shotgun metagenomic sequencing; library preparation, sequencing depth and coverage considerations, and read-based taxonomic classification using k-mer and alignment-based approaches.
Break – 00:30:00
Session 6 – 02:30:00 – Taxonomic classification and downstream analysis
Taxonomic classification vs profiling, tools, diversity analysis, differential abundance testing, and result visualization strategies.
Session 7 – 02:30:00 – Metagenomics assembly
Principles of de novo metagenomic assembly; assembly algorithms, coverage and community complexity challenges, quality metrics, and tools for short- and long-read data.
Break – 00:30:00
Session 8 – 02:30:00 – Draft genomes (bins) and Metagenome-Assembled Genomes (MAGs)
Binning strategies for recovering individual genomes from assemblies; coverage- and composition-based approaches, and generation of draft genomes and MAGs.
Session 9 – 02:30:00 – MAG quality assessment, annotation and classificationEvaluation of MAG completeness and contamination; genome annotation, taxonomic classification, and adherence to quality standards (e.g., MIMAG) for reporting draft genomes.
Break – 00:30:00
Session 10 – 02:30:00 – FAIR data and reproducibility
Application of FAIR principles to metagenomic datasets; public repository deposition, metadata standards, workflow managers, and containerization for reproducible analyses.
Frequently asked questions
Everything you need to know about the product and billing.
When will I receive instructions on how to join?
You’ll receive an email on the Friday before the course begins, with full instructions on how to join via Zoom. Please ensure you have Zoom installed in advance.
Do I need administrator rights on my computer?
I’m attending the course live — will I also get access to the session recordings?
I can’t attend every live session — can I join some sessions live and catch up on others later?
I’m in a different time zone and plan to follow the course via recordings. When will these be available?
I can’t attend live — how can I ask questions?
Will I receive a certificate?
When will I receive instructions on how to join?
You’ll receive an email on the Friday before the course begins, with full instructions on how to join via Zoom. Please ensure you have Zoom installed in advance.
Do I need administrator rights on my computer?
I’m attending the course live — will I also get access to the session recordings?
I can’t attend every live session — can I join some sessions live and catch up on others later?
I’m in a different time zone and plan to follow the course via recordings. When will these be available?
I can’t attend live — how can I ask questions?
Will I receive a certificate?
Still have questions?
Can’t find the answer you’re looking for? Please chat to our friendly team.








