Undergraduates & Masters
The Centre for Genomic Regulation (CRG) aims to provide highly motivated undergraduate and master students the opportunity to conduct research at the CRG. The goal is to encourage students (from all nationalities) in the pursuit of a scientific career giving them the prospect to get experience in an international laboratory while improving their skills.
The CRG is a center of excellence with international teams representing a broad range of disciplines, with first class core technologies to support the research projects, a wide range of seminars given by high-profile invited speakers, and courses on complementary and transferable skills integrated with the training programme.
We accept applications throughout the year for any type of internship with a learning agreement with your university. Have a look at our labs and research programmes and contact the Group Leader of your choice directly with the following documents attached:
- Motivation letter
- CV
- Reference letter
- University transcripts
Acceptance will depend on the capacity of the research group and the ongoing projects.
We host online events and workshops to inform you about various opportunities available at the CRG and guide you on how to search for PhD positions. If you are keen to learn more, please check HERE.

Fellowships available for the academic year 2026/2027:
We are pleased to offer 10 fellowship opportunities across different research projects at the CRG. Each project will remain open until the right candidate is found, giving motivated students the chance to join our vibrant research community and contribute to cutting-edge science.
Fellowships conditions
- Stipend: 700€/month/gross – up to 4 months
- Travel support: Return ticket (up to 1000€/non European flights / up to 400€ for European flights)
- Eligibility: The fellowship can only be given to new recruits, not students already at the CRG
- Timing: The fellowship needs to be given within the 2026/2027 academic year
Application Procedure:
Read carefully the projects below and, if interested, contact the project supervisor and group leader. Please include the following documents in your application:
- Motivation letter
- Curriculum Vitae (CV)
- Reference letter
- University transcripts
Proteomics - Sabido Core technologies
PROJECT DESCRIPTION - Development of robust LC-MS/MS workflows for proteoform characterization in liquid biopsies
Proteomics | LC-MS/MS | Mass spectrometry | Method development | Sample preparation | Liquid biological samples | Analytical workflow | Protein analysis
Mass spectrometry-based proteomics has become an indispensable tool for the comprehensive characterization of proteins in biological systems, providing valuable insights into cellular processes, disease mechanisms, and biomarker discovery. However, obtaining high-quality and reproducible proteomic data from liquid biological samples remains technically challenging due to sample complexity and variability. Robust sample preparation and optimized analytical workflows are therefore essential to maximize proteome coverage, reproducibility, and quantitative performance.
The aim of this project is to develop and optimize LC-MS/MS-based workflows for the proteomic analysis of liquid biopsies. The student will evaluate key steps in the experimental workflow, including sample preparation, protein digestion, peptide clean-up, and analytical performance, with the objective of establishing a robust and reproducible protocol suitable to perform proteoform characterization in liquid biopsies. Workflow performance will be assessed using quality metrics such as protein and peptide identification rates, reproducibility, and quantitative consistency.
Throughout the project, the student will gain practical experience in modern proteomics methodologies, including sample preparation, liquid chromatography-tandem mass spectrometry (LC-MS/MS), experimental design, quality control, and data analysis. The optimized workflow developed during this project will contribute to improving proteomic analyses performed within the CRG Proteomics Unit and will provide a valuable methodological framework for future biomedical research projects.
WHO ARE WE LOOKING FOR?
The project is suitable for a Master's student with a background in Biochemistry, Biotechnology, Molecular Biology, Biomedical Sciences, or a related field. Basic laboratory experience in molecular biology or biochemistry is desirable, while prior experience in proteomics or mass spectrometry is not required. The student should be motivated to learn new experimental techniques, have good organizational and analytical skills, and be interested in applying proteomics technologies to biomedical research.
HOW TO APPLY
Send applications documents to the PI and Hiba Salim - HERE
Regulatory genomics and diabetes - Ferrer Lab
PROJECT DESCRIPTION - CRISPR interference (CRISPRi) screening platform in human pancreatic beta cells
Type 1 Diabetes | CRISPR interference screening | CROP-seq | Islet cell therapy | Functional genomics
Type 1 diabetes (T1D) is an autoimmune disease characterised by the progressive destruction of insulin-producing pancreatic beta cells, leading to a lifelong dependence on exogenous insulin. Transplantation of cadaveric islets can restore glycaemic control, but the scarce supply has limited the widespread use of this therapy. Recent studies have enabled the efficient differentiation of human pluripotent stem cells into insulin-producing islet-like clusters (SC-islets), offering a scalable alternative source of cells for cell therapy, as reflected in ongoing clinical trials. Despite this progress, current SC-islets still show metabolic limitations and require immunosuppression, highlighting the need to further optimise beta-cell function. Machine learning approaches can be used to improve SC-islets. It is possible to build a virtual beta cell platform, able to predict how defined perturbations affect beta cell function, guiding the rational optimisation of SC-islets for T1D cell therapies.
To enable robust virtual cell models, it is critically important to generate genetic perturbation data in relevant cell types, which can be used to train the model. This project aims to establish the feasibility of a pooled CRISPR interference (CRISPRi) screening platform in human pancreatic beta cells. The screen will target genes central to beta cell identity, function, and stress resilience, including regulatory network hubs, transcription factors, genetic effectors of T1D and T2D risk, as well as genes that form part of pathways relevant to T1D pathophysiology and graft failure. The screening will be integrated with single-cell RNA sequencing using the CROP-seq approach, which allows simultaneous detection of transcriptional state and the corresponding genetic perturbation at the single-cell level. This strategy will set the basis for a virtual cell model capable of predicting beta-cell transcriptional responses to perturbations.
During this four-month project, the student will be trained in and will perform the key experimental steps required to generate the screening dataset: cloning of guide RNA constructs into CRISPRi backbone, production and titration of lentiviral vectors, transduction of human pancreatic beta cells, and generation of single-cell sequencing libraries. The resulting single-cell data will be analysed by a bioinformatics postdoctoral researcher in the laboratory, and the student will follow this analytical process to gain an understanding of the computational pipeline.
This project will provide training in functional genomics and single-cell technologies, while generating proof-of-concept data to establish the feasibility of the CROP-seq screening that will ultimately feed into the virtual beta cell model for optimising SC-islet-based T1D therapies.
WHO ARE WE LOOKING FOR?
The ideal candidate for this project is an undergraduate student with a background in biotechnology, molecular biology, genetics, or a related life sciences discipline. A basic understanding of molecular cloning, cell culture, and general genetics principles is desirable, though not strictly required, as full training will be provided. More important are genuine curiosity and motivation to learn hands-on wet lab techniques, including molecular cloning, lentiviral production, and cell culture, alongside a willingness to engage with the basics of dry lab work, such as understanding how single-cell sequencing data is processed and analysed. Strong attention to detail, reliability in following experimental protocols, and good communication skills are also valuable, given the collaborative nature of the project.
HOW TO APPLY
Send applications documents to the PI and Chiara Simoni - HERE
Single cell genomics and evolution - Sebé-Pedrós Lab
PROJECT DESCRIPTION - Comparing large-scale integration approaches for genome-wide expression in developmental trajectories
Single cell transcriptomics | Gene regulation | Genomics | Development | Computational biology | Systems biology
Single cell technologies have transformed our understanding of core developmental processes. Classic questions involving the nature of cell differentiation, and the extent of its conservation across animal lineages, can now be revisited with unprecedented cellular resolution by performing time-series single cell transcriptomics throughout development. To properly quantify and assess the conservation of cellular trajectories, we need to integrate (stitch) temporal atlases to relate cellular identities from consecutive stages in a rigorous way. To date, multiple approaches have been tested, but they have not been compared systematically. Here we propose a bioinformatics project to implement and compare three stitching approaches (graph-based methods using k-nearest-neighbour co-embedding, tree-based methods, and modality-specific modelling of cell-to-cell transitions) throughout animal development (one or more chordate species, like amphioxus, zebrafish, or mouse). The candidate will gain access to state-of-the-art methods and resources for analysing high-throughput single cell transcriptomics data, and their contributions will pave the way for a standardised method in large-scale comparative studies.
WHO ARE WE LOOKING FOR?
We are looking for a Life Sciences BSc/MSc with an interest in gene regulation and computational biology. Ideally, the student should have a background in development, genomics, bioinformatics, and computational biology. They will require basic skills in the command line (Bash/Linux) and R, as well as basic knowledge of Python. Prior experience with R programming and/or working in a high-performance cluster (HPC) will be highly valued. Organisation skills will also be valued.
HOW TO APPLY
Send applications documents to the PI and Alberto Pérez Posada - HERE
Systems and Synthetic Biology - Latorre Lab
PROJECT DESCRIPTION - Evolutionary CpG Spatial Organisation and Mammalian Longevity
Comparative genomics | CpG depletion | DNA methylation | Aging evolution | Bioinformatics
DNA methylation leaves a long-term evolutionary footprint on the genome sequence. Methylated cytosines at CpG dinucleotides are especially prone to mutation, contributing to CpG depletion over evolutionary time, while many regulatory regions remain comparatively CpG-rich. Previous comparative studies have linked CpG density in selected genomic regions to species lifespan and age at maturity. However, it remains unclear whether the spatial arrangement of CpGs contains information about longevity beyond their overall abundance.
This four-month project will develop a reproducible comparative-genomics workflow to measure distances between adjacent CpG sites across approximately 30-50 mammalian genomes. The student will construct empirical inter-CpG gap distributions and multiscale survival curves, then extract robust spatial descriptors and compare them with established measures. These genomic features will be integrated with publicly available lifespan, body-mass and maturity data. Phylogenetically controlled statistical models will test whether CpG spatial organisation explains variation in maximum lifespan or body-mass-adjusted longevity after accounting for shared ancestry and major technical covariates.
Mammals are used as the primary proof-of-concept because they combine broad lifespan variation with relatively comparable methylation biology, high-quality genome assemblies and well-curated life history information. If time permits, the pipeline will be tested on a small panel of non-mammalian vertebrates. The project will generate a documented analysis workflow, a quality-controlled comparative dataset and interpretable results, including the scientifically informative possibility that spatial descriptors do not outperform simpler composition metrics.
WHO ARE WE LOOKING FOR?
We are looking for a BSc or MSc student interested in bioinformatics and computational biology, with an enthusiasm for genomics, evolution, or ageing research. This is a computational (bioinformatics) project involving the analysis of large-scale genomic datasets. Suitable backgrounds include bioinformatics, biology, biotechnology, genetics, computer science, or another quantitative discipline.
Basic programming experience in Python or R is desirable. Familiarity with Linux, command-line tools, Git, statistics, or genomic file formats would be an advantage, but is not essential. More important are curiosity, careful data handling, a willingness to learn, clear documentation habits, and the ability to work independently while seeking guidance when needed.
The project is well suited to either a life-science student who wants to strengthen computational and bioinformatics skills or a quantitative student who wants to gain experience in biological data analysis. Prior knowledge of DNA methylation or phylogenetic comparative methods is welcome but not required; the relevant biological concepts, bioinformatics tools, and analytical methods will be introduced during the project.
HOW TO APPLY
Send applications documents to the PI and Eric Macwan - HERE
Reprogramming and Regeneration - Cosma Lab
PROJECT DESCRIPTION - Selective depletion of TOP2A or TOP2B alters chromatin folding in human HCT116 cells
ORCA | Chromatin folding | DNA topology | Topoisomerases | Single-cell imaging | Quantitative analysis
The eukaryotic genome is folded into a highly organised three-dimensional (3D) architecture that is essential for gene regulation and genome stability. Chromatin loops formed by cohesin and CTCF are major components of this organisation, but it remains unclear how DNA topology, including torsional stress and supercoiling, contributes to their formation and stability.
DNA topoisomerases regulate the topological stress generated during transcription and replication. Preliminary imaging and genomic analyses from our laboratory suggest that perturbation of TOP2A or TOP2B affects DNA compaction and chromatin organisation. However, whether these changes reflect alterations in the 3D folding of individual genomic regions remains unknown.
The objective of this project is to determine whether selective depletion of TOP2A or TOP2B alters chromatin folding at a defined genomic locus in human HCT116 cells. Optical Reconstruction of Chromatin Architecture (ORCA) will be used to address this question. ORCA combines sequential rounds of fluorescence imaging to visualise consecutive regions across a genomic locus and reconstruct their spatial organisation in individual cells.
Control cells will be compared with cells depleted of TOP2A or TOP2B. The analysis will quantify pairwise distances and contact frequencies between selected genomic regions, local chromatin compaction and cell-to-cell variability. This comparison will reveal whether the two topoisomerase II isoforms contribute similarly or differently to local chromatin organisation and will provide insight into how DNA topology shapes genome folding at the single-cell level.
The target locus, probe library, imaging workflow and core analysis pipeline will be established before the fellowship begins, allowing the project to focus on the biological comparison and on the generation, analysis and interpretation of a well-defined ORCA dataset.
WHO ARE WE LOOKING FOR?
The project is suitable for a motivated undergraduate student interested in molecular and cell biology, microscopy or quantitative biology. The student does not need to have experience in all of these areas. No previous knowledge of ORCA, advanced microscopy or programming is required.
The most important qualities are curiosity, reliability, careful working practices and a willingness to learn. Basic knowledge of molecular and cell biology would be helpful, while previous exposure to microscopy, image analysis or Python would be an advantage but is not essential.
The balance between experimental and computational work can be adapted to the student's interests. A student more interested in experimental biology may focus primarily on cell culture, sample preparation, imaging and biological interpretation. A student with stronger quantitative interests may focus more on image analysis, data visualisation and Python-based analysis. If the student progresses rapidly and is particularly comfortable with computational analysis, the project may be extended to include introductory integration of ORCA measurements with Hi-C data using an established analysis and modelling pipeline. This additional component would remain optional and would not be required for successful completion of the fellowship.
Regardless of their starting profile, the student will receive structured, step-by-step training and close supervision throughout the fellowship. They will not be expected to establish ORCA independently or design the probe library, as these elements will already be available at the beginning of the project.
HOW TO APPLY
Send applications documents to the PI and Mégane Da Mota - HERE
Cell fate decoding and engineering - Lin Lab
PROJECT DESCRIPTION - Understanding neuronal diversity
Cell fate engineering | Transcription factors | Molecular cloning | Lentiviral vectors | Mammalian cell culture
Cell fate engineering holds enormous potential for both basic research and translational applications, yet current in vitro models capture only a small fraction of the neuronal diversity present in the human brain. While thousands of neuronal subtypes arise in vivo through the coordinated action of regional patterning cues and dynamic transcriptional programs, many existing engineering approaches rely on the overexpression of a limited set of transcription factors from pluripotency to directly generate specific neuronal identities. How combinations of transcription factors interact with distinct progenitor states to generate neuronal diversity is currently underexplored.
This project aims to address these limitations through a systematic transcription factor perturbation framework guided by gene regulatory network inference from human brain atlases. Candidate transcription factors will be assembled into optimized delivery vectors and introduced into regionally specified progenitor cells to investigate how developmental context shapes transcription factor-driven fate conversion. We will explore combinatorial transcription factor perturbations and seek to uncover multiple neuronal identities simultaneously and define how progenitor state and transcription factor interactions cooperate to generate cellular diversity.
Overall, this work aims to establish a scalable framework for understanding and engineering neuronal diversity, providing new insights into the developmental principles that govern cell fate specification and contributing to the generation of more physiologically relevant human neural models.
WHO ARE WE LOOKING FOR?
The project is suitable for students from a range of backgrounds, as we will provide training from the ground up. More important than prior experience is motivation, curiosity, and a willingness to learn and actively integrate into the group. Previous experience in molecular cloning would be an advantage, but it is not required.
HOW TO APPLY
Send applications documents to the Albert Blanch - HERE
Probabilistic machine learning and genomics - Dias & Frazer Lab
PROJECT DESCRIPTION - Folding versus function: disentangling the molecular mechanisms of missense variants in gene-trait associations
Protein language models | Missense variants | Protein stability | Protein folding | Human genetics | UK Biobank | Variant effect prediction | Computational biology | Machine learning | Protein biophysics
Genome sequencing has grown exponentially, but determining the phenotypic consequences of missense variants remains a key challenge in human genetics. Deep-learning unsupervised methods, such as protein language models (pLM), that learn the distribution of sequence variation across organisms have emerged as promising tools for scoring variant effects. In gene-trait association studies, regression tests using pLMs find ∼50% more associations than burden tests (Jang et al., Cell Genomics 2026) and can be run directly from precomputed summary statistics (Dinh et al., Nature Methods 2026). However, pLM are trained on evolutionary data and therefore often fall short of attributing effects to specific biophysical mechanisms.We propose to integrate protein language models with folding energetics to disentangle effects on stability and on functions beyond folding, such as allostery, ligand/partner binding and downstream interactions. We will quantify and separate mutational effects on the energetics of protein folding versus natural selection by computing protein “dark energy”, defined as the difference between physical folding free energies and evolutionary free energies derived from a pLM (Galpern et al., PNAS 2026). We will apply this biophysical decomposition to study the genotype-phenotype summary statistics from the UK Biobank as provided by Genebass (Karczewski et al., Cell Genomics 2022), focusing specifically on genes associated with blood biochemistry biomarkers. We will relate the predicted folding-stability and dark energy changes for each missense variant to the effect sizes and contrast those regressions against the gene's putative loss-of-function (pLoF) effect. We will disentangle variants destabilizing the encoded protein from others that preserve protein stability but affect specific chemical activities. Together, this work will pave the way to a scalable framework to move beyond statistical gene–trait associations toward a residue-level, mechanistic understanding of how missense variation perturbs protein biology in human genetics.
WHO ARE WE LOOKING FOR?
We are looking for a motivated master student with a background in bioinformatics, computational biology, computer science, physics, mathematics, or a related quantitative discipline. Previous programming experience in Python is expected. Familiarity with molecular biology, genetics, statistics, or machine learning is desirable but not required. The student should be curious and interested in applying computational methods to biological questions.
HOW TO APPLY
Send applications documents to the PI and Ezequiel Galpern - HERE
Systems & Synthetic Biology - MARTIN Lab
PROJECT DESCRIPTION - Understanding the effect of environmental changes on a genotype-phenotype map
Evolution | Genotype-phenotype maps | Fitness seascapes
Models of evolution using genotype-phenotype maps and the resulting fitness landscapes allow for the accurate treatment of genetic sequences and the high-dimensional structure of mutations. Improvements in both computational resources and experimental methods make it possible to study increasingly large landscapes, but there are still many unanswered questions about what structures can and do exist in real genotype spaces, and how an individual's changing environment influences both its phenotype and fitness. This project will use computational tools to understand the effect of a changing environment on a genotype-phenotype map, and on the subsequent effect that a changing genotype-phenotype map has on evolutionary simulations.
WHO ARE WE LOOKING FOR?
As an interdisciplinary field, this project will be most interesting to a student either:
- With a physics/computer science background and an interest in learning more about biology and genetics, or
- With a biology/bioinformatics background with an interest in learning more about programming and statistics.
A strong motivation to learn about new topics is essential.
HOW TO APPLY
Send applications documents to the PI and Kye Hunter - HERE
Computational Biology and Health Genomics - Bernardo Rodríguez Martín Lab
PROJECT DESCRIPTION - Mapping regulatory sequence motifs in repetitive structural variants to interpret their impact on gene regulation
Structural variants | Long-read sequencing | Gene regulation | Mobile elements | Tandem repeats | Transcription factors | RNA binding proteins
Understanding how disease-associated variants impact gene function, and pinpointing the underlying causal variant, remains challenging despite decades of association studies like GWAS and expression QTLs. Often overlooked as the potential mechanism driving these associations are structural variants (SVs). Although SVs represent the largest source of genetic variation between human genomes by megabase pairs (Collins and Talkowski 2025), and are more likely than single nucleotide variants to be causal (Chiang et al. 2017; Bai et al. 2026), their repetitive nature makes them difficult to detect and interpret, leaving them largely understudied.
Excitingly, long-read sequencing technologies are helping bridge the gap. We recently produced the largest map of human SVs to date by analyzing 1019 long-read human genomes, detecting ~two-fold more SVs detected per genome compared to short reads and fully resolving their sequences (Schloissnig et al. 2025). Crucially, by studying the previously inaccessible internal sequence we can gain insights into their functional effects. For example, CpGs in tandem repeats, a type of SV with variable lengths across individuals, can lead to hypermethylation resulting in abnormal gene silencing in disease (Hannan 2018). Similarly, mobile element insertions often contain splice and transcription factor binding sites, leading to aberrant isoforms and ectopic oncogene expression (Payer and Burns 2019; Diaz-Portal et al. 2026).
In this project, we propose to analyze their internal sequence as a source of evidence for interpreting SVs in disease-associated loci. We have recently annotated the sequences of SVs derived from 1019 long-read human genomes from diverse populations, and found extensive variability both between repeat copies (Schloissnig et al. 2025). The student will help us make sense of these layers of variation by mapping transcription factor and RNA-binding protein binding sites within these sequences (Grant et al. 2011), and explore predictive approaches, such as AlphaGenome (Avsec et al. 2026), to link the presence of motifs with gene expression variation.
Overall, the annotations will serve as a basis to interpret the functional effects of some of the most difficult-to-study variants in the genome. These results will identify SVs likely impacting the regulation of nearby genes, and allow us to propose an underlying molecular mechanism of disease-associated loci.
WHO ARE WE LOOKING FOR?
Skills: Python, Bash/command line
Knowledge: molecular biology, statistics, genomics
Background: bioinformatics, biotechnology, genomic sciences, or similar.
HOW TO APPLY
Send applications documents to the PI and Jesus Emiliano Sotelo - HERE
Evolutionary Processes Modeling - Weghorn Lab
PROJECT DESCRIPTION - How do genetic and evolutionary factors shape the immunogenic potential of cancer-associated mutations
Evolution | Neoantigen | Immunogenicity
Cancer develops through the accumulation of somatic mutations, some of which generate neoantigens that can be presented by human leukocyte antigen (HLA) molecules and recognized by T cells. These interactions may influence which mutations persist during tumour evolution, but the extent and determinants of immune-mediated selection remain unclear. Progress has been limited by the difficulty of modelling mutation probabilities, antigen presentation, and T-cell recognition with sufficient accuracy. This project will investigate how genetic and evolutionary factors shape the immunogenic potential of cancer-associated mutations. Using comparative sequence analysis, evolutionary simulations, and context-aware computational models, we will examine how variation in coding sequences affects the likelihood that somatic mutations generate peptides with the potential for immune recognition. The work will provide a broader understanding of the relationship between genome evolution, somatic mutation, and tumour immunogenicity, and may help clarify the factors that influence immune selection during cancer development.
WHO ARE WE LOOKING FOR?
A student with:
- An interest in computational biology, evolution, or cancer genomics.
- Experience in Python (or similar) and handling large datasets.
- Some familiarity with bioinformatics pipelines (preferred but not required).
HOW TO APPLY
Send applications documents to the PI, María Kelly and Francisco Javier Ordoñez - HERE
Contact
For any further questions, please contact
CRG Training & Academic Office
Centre de Regulació Genòmica
Dr. Aiguader, 88
PRBB Building
08003 Barcelona
training@crg.eu
