Research

The Pereira lab focuses on creating a broad, detailed overview of protein sequence, structure, and function data. 

The team combines this with machine learning and network analysis to predict, label, and map unexplored areas of the protein world, especially those relevant to medicine and biotechnology on a large scale. 

The goal is to develop a thorough map of unknown biological information hidden in large protein databases. This will help identify and understand proteins that could be part of undiscovered biological systems or functions, and prioritize them for further study and potential use in biotechnology.

Protein universe

Mapping the entire protein universe

The protein landscape currently mapped in the Protein Universe Atlas covers only ~10% of all proteins in UniProt, and includes only those proteins for which AlphaFold2 has been able to model in 3D with a high degree of accuracy. Billions of predicted proteins are still missing. 

We aim to extend this landscape to the entire publicly available protein catalogue, and develop methods for efficient interpretation of evolutionary relationships and automatic identification and classification of novel protein families and superfamilies within this ever-growing data set.

Protein

Identifying unknown proteins with impact on our lives

We use the landscape we've built to uncover unknown biological systems, explore their origins, and better understand their functions in biology.

From novel transmembrane signalling detectors and channels to novel defence mechanisms, virulence factors, or putative metabolic modulators--we want to provide a comprehensive, annotated map of all proteins of unknown function, giving priority to those that may have a direct impact on human health or a potential for novel biotechnological or biomedical applications.

Expanding GCsnap

GCsnap is a tool that makes it easy to compare conserved genomic contexts, while also incorporating functional and structural information. It works well for focused studies on specific gene families, but needs further development to efficiently handle large-scale analysis of millions of protein sequence groups.

We are working to enhance GCsnap by adding features that predict and annotate structural models for potential protein interactions, based on their genomic neighborhood. This will greatly improve its usefulness and real-world applications. Our goal is to turn GCsnap into a comprehensive, user-friendly toolkit that can infer biological functions by combining genomic context, predicted structures, functional insights, and potential impacts of gene loss on metabolism.

“The ultimate goal is to cover the entire universe of natural proteins, and to predict, annotate, and classify those proteins which may have a direct or indirect impact on human health.”
Joana Pereira
Joana Pereira
Group leader

Selected publications

Image
Protein universe

Uncovering new families and folds in the natural protein universe

Text

Durairaj J, Waterhouse AN, Mets T, Brodiazhenko T, Abdullah M, Studer G, Tauriello G, Akdel M, Andreeva A, Bateman A, Tenson T, Hauryliuk V, Schwede T*, Pereira J*.

Image
Protein structure

How Do I Get the Most Out of My Protein Sequence Using Bioinformatics Tools?

Text

Pereira J, Alva V.  

Image
GCsnap

GCsnap: Interactive Snapshots for The Comparison of Protein-Coding Genomic Contexts

Text

Pereira J.

Image
Barrel

Gram-Negative Outer Membrane Proteins with Multiple β-Barrel Domains

Text

Solan R*, Pereira J*,  Lupas AN, Kolodny R, Ben-Tal N.

Image
Main chain

A Distance Geometry-based Description of Protein Main-Chain Conformational Space

Text

Pereira J, Lamzin V.

For a full publication overview visit