P ProteinPRS Polygenic Prediction of Plasma Protein Levels

Sparse Genome-Wide Polygenic Prediction of plasma protein levels.

The ProteinPRS portal enables users to query, visualize, and download SNPBoost-derived PRS models for serum protein concentration.

These models, trained to capture both local cis and distant trans-regulatory mechanisms, identify the most informative variants for protein levels through multivariable regression within a boosting framework, aiding in the discovery of independent pQTL loci.

Furthermore, the models are instrumental for inferring genetically driven proteomic levels in independent datasets, thus supporting proteome-wide association studies (PWAS) by associating imputed protein levels with specific phenotypes.

2,923
proteins
52,705
samples
1.14 M
selected variants
Olink
Explore 3072

Methods

Data

Olink measurements for nearly 3,000 plasma proteins across more than 50,000 UK Biobank samples, predominantly of European ancestry.

Algorithm

SNPBoost produces sparse polygenic models from genome-wide data, accounting for the joint effect of several variants rather than marginal associations.

Interpretation

Cis variants may be pQTLs or their LD proxies. Distant variants can reflect transcription-factor effects or protein–protein interactions on concentration.

Limitations

Plasma concentration only — no tissue or cell specificity. Performance may drop in populations other than the training population.

Contact

About

We are actively refining PRS models to enhance protein level predictions and elucidate genetic loci implicated in regulatory mechanisms of protein expression.

Contact

For questions and collaboration inquiries, please contact Dr. Maj: carlo.maj@uni-marburg.de

Protein PRS Models

Genome-wide model R² against cis-model R² for every protein, coloured by model sparsity.

{{ scatterGrid }} {{ scatterFrame }}

All proteins

{{ rangeLabel }} · drag horizontally for all 17 fields
{{ pageLabel }}
Gene
test_n
sparsity
pearson_coefficent
SE
P_corr
R2 Covariate Model
R2 Full Model
P
P_full
n_cis_pqtl
max_cis_pqtl
rank_max_cis_pqtl
mean_pqtl_beta
mean_nopqtl_beta
percentage_pqtl
percentage_nopqtl
{{ row.gene }}
{{ row.test_n }}
{{ row.sparsity }}
{{ row.pearson }}
{{ row.se }}
{{ row.p_corr }}
{{ row.r2_cov }}
{{ row.r2_full }}
{{ row.p }}
{{ row.p_full }}
{{ row.n_cis }}
{{ row.max_cis }}
{{ row.rank_cis }}
{{ row.mean_pqtl }}
{{ row.mean_nopqtl }}
{{ row.pct_pqtl }}
{{ row.pct_nopqtl }}
No protein matches that filter.

{{ sel.gene }}

{{ sel.name }}

Download PRS model
R2 Full Model
{{ sel.r2 }}
pearson_coefficent
{{ sel.corr }}
sparsity
{{ sel.sparsity }}
n_cis_pqtl
{{ sel.n_cis }}
max_cis_pqtl
{{ sel.max_cis }}

Diagnostics

PRS scatter and box plots enlarge ⤢
{{ chartPrs }}
Genome-wide Manhattan Plot enlarge ⤢
{{ chartGw }}
Regional Manhattan Plot enlarge ⤢
{{ chartReg }}
{{ modalTitle }}
{{ modalChart }}

Model record

Field
Value
{{ m.key }}
{{ m.value }}

cis and genome-wide metrics

Field
Value
{{ m.key }}
{{ m.value }}

FAQ

1. What is ProteinPRS?

ProteinPRS is a web portal designed for the visualization, querying, and downloading of polygenic risk score (PRS) models for plasma protein levels.

These models are based on Olink data for nearly 3,000 proteins measured in over 50,000 samples from the UK Biobank.

The models are derived using the SNPBoost algorithm (SNPBoost), which efficiently processes high-dimensional omics data to generate optimized sparse polygenic models taking into account the joint effect of several variants.

2. Why are ProteinPRS models useful for researchers in molecular biology?

ProteinPRS models incorporate complete genome-wide data, capturing both local cis-regulatory effects and distant regulatory mechanisms. The selected variants in the cis-region may potentially be pQTL regulatory variants or their linkage disequilibrium (LD) proxies. The selected variants in distant regions can represent the effect of transcription factors and/or disclose protein-protein interaction effects affecting protein level concentration.

3. Why are ProteinPRS models useful for researchers investigating genetic associations with complex phenotypes?

Similar to transcriptome-wide association studies (TWAS), ProteinPRS models enable the inference of the genetically regulated component of protein expression in an independent dataset.

The imputed protein levels can then be associated with available phenotypes, as in a proteome-wide association study (PWAS).

This approach can be utilized for gene-based prioritization of significant loci identified by genome-wide association studies (GWAS).

4. How can I predict the genetically-driven component of protein expression in my dataset using ProteinPRS models?

The PRS models for each protein can be downloaded as scoring files compatible with the PLINK2 score function, including variant identifiers as rsid, effect alleles, and weights.

The scoring function can be applied to standard genetic data input formats (e.g., PLINK, binary PLINK, VCF, etc.).

Additional annotation files with genomic coordinates in hg19 and hg38 are also provided to adapt the scoring files with possible different variant annotations (e.g., chr:bp:ref:alt)

5. Have the methods and results of ProteinPRS models been published?

An abstract on SNPBoost-based protein level prediction has been submitted to ESHG 2024 (ESHG 2024), and a corresponding manuscript is under preparation.

However, different studies using the implemented boosting algorithm for classical phenotype prediction have already been published:

6. Are there alternative PRS models for protein level predictions?

With the growing availability of large-biobank data and intensive research efforts, several alternatives likely exist.

A major reference database for polygenic risk score prediction of molecular markers, including proteomics data, is OMICSPRED (OMICSPRED).

However, most models are typically based only on local cis-expression regulation and neglect distant trans-pQTL effects.

Thus ProteinPRS models by taking into account distant pQTL effects might improve protein prediction and also precision for GWAS prioritization (while cis-eQTL are expected to be correlated at in nearby genes, the trans-eQTL are expected to be independent across genes in a locus).

7. What are the current limitations of ProteinPRS models?

The major limitations are as follows:

  1. Protein expression is primarily environmentally driven; thus, the genetic prediction is limited and heterogeneous across genes. It is advisable to check the model prediction performance (e.g., Pearson correlation and p-value) for individual proteins since the genetic regulation can significantly influence protein concentration for some proteins, whereas for others, the effect may be negligible.
  2. The models are based on plasma level concentrations, not accounting for tissue and cell specificity of protein regulation.
  3. The models were trained on predominantly European sample data. Although gene-expression regulation may be less impacted by population-specific effects with respect to disease phenotypes, also ProteinPRS models, as all PRS models, might have reduced performances when tested on different populations.