Research
The Pereira lab focuses on creating a broad, detailed overview of protein sequence, structure, and function data.
The team combines this with machine learning and network analysis to predict, label, and map unexplored areas of the protein world, especially those relevant to medicine and biotechnology on a large scale.
The goal is to develop a thorough map of unknown biological information hidden in large protein databases. This will help identify and understand proteins that could be part of undiscovered biological systems or functions, and prioritize them for further study and potential use in biotechnology.
Mapping the entire protein universe
The protein landscape currently mapped in the Protein Universe Atlas covers only ~10% of all proteins in UniProt, and includes only those proteins for which AlphaFold2 has been able to model in 3D with a high degree of accuracy. Billions of predicted proteins are still missing.
We aim to extend this landscape to the entire publicly available protein catalogue, and develop methods for efficient interpretation of evolutionary relationships and automatic identification and classification of novel protein families and superfamilies within this ever-growing data set.
Identifying unknown proteins with impact on our lives
We use the landscape we've built to uncover unknown biological systems, explore their origins, and better understand their functions in biology.
From novel transmembrane signalling detectors and channels to novel defence mechanisms, virulence factors, or putative metabolic modulators--we want to provide a comprehensive, annotated map of all proteins of unknown function, giving priority to those that may have a direct impact on human health or a potential for novel biotechnological or biomedical applications.
Expanding GCsnap
GCsnap is a tool that makes it easy to compare conserved genomic contexts, while also incorporating functional and structural information. It works well for focused studies on specific gene families, but needs further development to efficiently handle large-scale analysis of millions of protein sequence groups.
We are working to enhance GCsnap by adding features that predict and annotate structural models for potential protein interactions, based on their genomic neighborhood. This will greatly improve its usefulness and real-world applications. Our goal is to turn GCsnap into a comprehensive, user-friendly toolkit that can infer biological functions by combining genomic context, predicted structures, functional insights, and potential impacts of gene loss on metabolism.
Selected publications
Uncovering new families and folds in the natural protein universe
Durairaj J, Waterhouse AN, Mets T, Brodiazhenko T, Abdullah M, Studer G, Tauriello G, Akdel M, Andreeva A, Bateman A, Tenson T, Hauryliuk V, Schwede T*, Pereira J*.
How Do I Get the Most Out of My Protein Sequence Using Bioinformatics Tools?
Pereira J, Alva V.
GCsnap: Interactive Snapshots for The Comparison of Protein-Coding Genomic Contexts
Pereira J.
Gram-Negative Outer Membrane Proteins with Multiple β-Barrel Domains
Solan R*, Pereira J*, Lupas AN, Kolodny R, Ben-Tal N.
A Distance Geometry-based Description of Protein Main-Chain Conformational Space
Pereira J, Lamzin V.