£400Registration Fee
Register Now- Overview
- Instructors
- Schedule
Course Description
This course will cover the metagenomics data analysis workflow from data generation to downstream analysis. Participants will explore the tools to manage, share, analyze and interpret metagenomics data. The content will include issues of data quality control and how to process the data until publication.
Both targeted marker-gene (e.g., 16S) and whole-genome shotgun (WGS) approaches will be covered. Discussions will also explore considerations when selecting the sequencing options (long or short reads), exploring targeted data to OTUs or ASVs, assembling and binning metagenomics data, available tools, and the downstream analyses that can be carried out.
What You’ll Learn
- Define the core concepts, terminology, and scope of metagenomics, and distinguish it from related fields such as genomics, culturomics, and single-cell sequencing.
- Explain the biological and ecological rationale for studying microbial communities in situ, including the concept of the “unculturable majority.”
- Describe key considerations in experimental design for metagenomic studies, including replication, controls (negative/positive), and confounding variables.
- Recognize appropriate sample collection, storage, and DNA/RNA extraction methods for different environments (e.g., soil, gut, water) and their impact on downstream results.
- Interpret and compare the two main methodological approaches (targeted vs shotgun) metagenomics including their respective advantages, limitations, costs, and appropriate use cases.
- Identify relevant sequencing platforms (e.g., Illumina, Oxford Nanopore, PacBio) and how platform choice influences experimental design and analysis strategy.
- Execute quality control and pre-processing of raw sequencing data, including adapter trimming, quality filtering, host/contaminant removal, and assessment of sequencing depth and coverage.
- Explain the principles behind key computational steps such as read assembly, binning, taxonomic classification, and functional annotation.
- Evaluate the outputs of metagenomic pipelines critically, including assessment of assembly quality, bin completeness/contamination and classification confidence.
- Describe downstream analysis options, including diversity analysis (alpha/beta diversity), differential abundance testing, functional pathway analysis, and community comparison across conditions.
- Apply appropriate statistical methods and visualization techniques to interpret and communicate metagenomic results.
- Apply FAIR data principles and other good practices in bioinformatics to metagenomic datasets, including proper deposition in public repositories.
.
Course Format
Flexible Learning Structure
Learn through a carefully structured mix of lecture recordings and guided exercises that you can pause, revisit, and complete at your own pace—ideal for busy professionals or those balancing multiple commitments.
Access Anytime, Anywhere
All course content is available on-demand, making it accessible across all time zones without the need to attend live sessions or adjust your schedule.
Independent Exploration with Support
Engage deeply with course topics through self-directed study, with the option to reach out to instructors via email for clarification or deeper discussion.
Comprehensive Learning Resources
Gain full access to the same high-quality materials provided in live sessions, including code, datasets, and presentation slides—all available to download and keep. Please note recordings can only be streamed.
Work With Your Own Data, On Your Terms
Apply what you learn directly to your own data projects as you go, allowing for a personalized and immediately practical learning experience.
Continued Guidance and Resource Access
Receive 30 days of post-enrolment email support and unrestricted access to all session recordings during that time, so you can review and reinforce your learning as needed.
Who Should Attend / Intended Audiences
This course is designed for PhD students, postdoctoral researchers, and other life scientists from academia and industry who are working in metagenomics and are in the early stages of data preparation or analysis.
To get the most from the course, participants should have knowledge equivalent to the learning outcomes of First Steps with R in Life Sciences and the UNIX Fundamentals e-learning course.
Specifically, participants should be familiar with Rstudio, how to install a library, matrix and data frame manipulation, import and export data from text files. In case of doubt, evaluate your R skills with this quiz, before registering.
Equipment and Software requirements
You will be required to create a GitHub account at: https://github.com/signup. During the course, we will use the free tier of GitHub Codespaces in order to count with a standardized working environment for all the participants.
⚠ WARNING! Do not start using the Codespaces until the day of the course since it has limited computing hours and storage space, and they can be run out pretty fast.
A working webcam is recommended to support interactive elements of the course. We encourage participants to keep their cameras on during live Zoom sessions to foster a more engaging and collaborative environment.
While not essential, using a large monitor—or ideally a dual-monitor setup—can significantly enhance your learning experience by allowing you to view course materials and work in R simultaneously.
All necessary packages will be introduced and installed during the workshop. A comprehensive list of required packages will also be shared with participants ahead of the course to allow for optional pre-installation.
Dr. Jeferyd Yepes
Dr. Jeferyd Yepes
Jeferyd is a bioinformatician and computational biologist specialising in metagenomics, microbial genomics, reproducible bioinformatics workflows, and artificial intelligence for protein function prediction. His research combines large-scale metagenomic analysis with modern machine learning approaches to understand complex microbial communities and their functional potential.
Jeferyd completed his PhD in Bioinformatics at the University of Fribourg and the Swiss Institute of Bioinformatics (SIB), supported by a Swiss Government Excellence Scholarship. His doctoral research focused on understanding rice straw degradation using metagenomics, with particular emphasis on reconstructing Metagenome-Assembled Genomes (MAGs) to identify microorganisms with lignocellulose-degrading potential.
Alongside his metagenomics research, Jeferyd works with protein language models (pLMs), incorporating structural information and fine-tuning machine learning models to improve automated protein function prediction from microbiome data. His research has received international recognition, including selection as an SIB Remarkable Output in 2024 and the Best Poster in Bioinformatics and Systems Biology at the Life Sciences Switzerland (LS2) Annual Meeting 2025.
Education & Career
- PhD in Bioinformatics, University of Fribourg, Switzerland (2026)
- Bioinformatics Researcher, Swiss Institute of Bioinformatics (SIB), Switzerland
- Master’s degree in Engineering, University of Antioquia, Colombia
- Former Laboratory Analyst at Iluma Alliance and the University of Antioquia, Colombia
- Recipient of a Swiss Government Excellence Scholarship for doctoral research
Research Focus
Jeferyd’s research centres on developing and applying computational approaches for analysing complex microbial communities and extracting functional information from large-scale metagenomic datasets. His research interests include:
- Metagenomics and microbial community analysis
- Metagenome-Assembled Genome (MAG) reconstruction
- Microbial genome recovery, quality assessment, and taxonomic annotation
- Bioinformatics pipeline development and benchmarking
- Nextflow and reproducible computational workflows
- Protein language models and machine learning
- Protein function prediction from microbiome data
- Lignocellulose degradation and microbial biotechnology
Main Achievements
During his PhD research, Jeferyd:
- Developed a suite of bioinformatics tools supporting metagenomics pipeline selection, benchmarking, development, and reproducible analysis
- Performed comprehensive metagenomic analyses, including MAG reconstruction, identifying microorganisms with substantial lignocellulose-degrading potential during rice straw decomposition
- Applied and fine-tuned protein language models to improve prediction of lignocellulose-degrading enzyme functions, outperforming traditional sequence-based approaches
- Developed resources designed to make sophisticated metagenomics workflows more accessible and reproducible for researchers
Awards
- SIB Remarkable Output 2024 — Swiss Institute of Bioinformatics
- Best Poster in Bioinformatics and Systems Biology — Life Sciences Switzerland (LS2) Annual Meeting 2025
- Swiss Government Excellence Scholarship (2022) — Federal Commission for Scholarships for Foreign Students
Teaching & Training
Jeferyd has taught bioinformatics and computational biology through the Swiss Institute of Bioinformatics, with particular expertise in metagenomics, reproducible workflows, Nextflow, and scientific data visualisation. Previous courses include:
- Nextflow in Action: Build Smarter, Faster, Reproducible Pipelines (2026)
- Introduction to Metagenomics Data Analysis of Microbial Communities (2024, 2025, 2026)
- Interactive Visualization with Python Fribourg (2024, 2025)
- Practical training in reproducible bioinformatics workflows and computational analysis
Tools Developed
- 2Pipe — A decision-support tool designed to help researchers select appropriate pipelines for Metagenome-Assembled Genome reconstruction
- MAGFlow and BIgMAG integration — Tools for visualising metagenome quality metrics and taxonomic annotations
- TaxoFlow — A step-by-step framework and tutorial for constructing Nextflow pipelines for metagenomic taxonomic classification
Links
Session 1 – 02:30:00 – Introduction to Metagenomics, goals and opportunities
Overview of metagenomics concepts and history; culture-independent approaches, study design considerations, and applications across environmental, clinical, and biotechnological fields; introduction to sequencing technologies, short and long reads.
Break – 00:30:00
Session 2 – 02:30:00 – Quality Control and Contamination Assessment
Quality assessment of sequencing data; adapter removal, filtering strategies, contamination detection, and best practices for genomic and transcriptomic analyses.
Session 3 – 02:30:00 – Targeted sequencing I (Introduction and theory)
Principles of amplicon-based sequencing; marker gene selection (16S, ITS), primer design, PCR amplification biases, and library preparation strategies for targeted metagenomics; OTUs vs ASVs.
Break – 00:30:00
Session 4 – 02:30:00 – Targeted sequencing II (Downstream analyses)
Processing of amplicon sequencing data; OTU/ASV inference, taxonomic assignment, diversity metrics, and statistical approaches for microbial community comparison.
Session 5 – 02:30:00 – Whole genome sequencing and taxonomic classification
Principles of shotgun metagenomic sequencing; library preparation, sequencing depth and coverage considerations, and read-based taxonomic classification using k-mer and alignment-based approaches.
Break – 00:30:00
Session 6 – 02:30:00 – Taxonomic classification and downstream analysis
Taxonomic classification vs profiling, tools, diversity analysis, differential abundance testing, and result visualization strategies.
Session 7 – 02:30:00 – Metagenomics assembly
Principles of de novo metagenomic assembly; assembly algorithms, coverage and community complexity challenges, quality metrics, and tools for short- and long-read data.
Break – 00:30:00
Session 8 – 02:30:00 – Draft genomes (bins) and Metagenome-Assembled Genomes (MAGs)
Binning strategies for recovering individual genomes from assemblies; coverage- and composition-based approaches, and generation of draft genomes and MAGs.
Session 9 – 02:30:00 – MAG quality assessment, annotation and classificationEvaluation of MAG completeness and contamination; genome annotation, taxonomic classification, and adherence to quality standards (e.g., MIMAG) for reporting draft genomes.
Break – 00:30:00
Session 10 – 02:30:00 – FAIR data and reproducibility
Application of FAIR principles to metagenomic datasets; public repository deposition, metadata standards, workflow managers, and containerization for reproducible analyses.
Frequently asked questions
Everything you need to know about the product and billing.
When will I receive instructions on how to join?
You’ll receive an email on the Friday before the course begins, with full instructions on how to join via Zoom. Please ensure you have Zoom installed in advance.
Do I need administrator rights on my computer?
I’m attending the course live — will I also get access to the session recordings?
I can’t attend every live session — can I join some sessions live and catch up on others later?
I’m in a different time zone and plan to follow the course via recordings. When will these be available?
I can’t attend live — how can I ask questions?
Will I receive a certificate?
When will I receive instructions on how to join?
You’ll receive an email on the Friday before the course begins, with full instructions on how to join via Zoom. Please ensure you have Zoom installed in advance.
Do I need administrator rights on my computer?
I’m attending the course live — will I also get access to the session recordings?
I can’t attend every live session — can I join some sessions live and catch up on others later?
I’m in a different time zone and plan to follow the course via recordings. When will these be available?
I can’t attend live — how can I ask questions?
Will I receive a certificate?
Still have questions?
Can’t find the answer you’re looking for? Please chat to our friendly team.








