Feed 0% source
Chemistry AI-generated

Beyond the classical plasma secretome: genetic architecture and disease associations of the expanded human plasma proteome in 13,445 Europeans

Generated by a local model (nvidia/Gemma-4-26B-A4B-NVFP4) from a scientific paper, claim-checked against the full text. Provenance is open by design.

Beyond the Classical Secretome: Unlocking the Hidden Proteome

Scientists use advanced tools to measure over 7,000 proteins in the blood instead of the usual 5,000. This expansion allows them to find new genetic links to diseases. It also discovers that many proteins previously ignored—specifically those typically found inside cells—play vital roles in human health.

For years, large-scale proteomics (the study of the entire set of proteins expressed by a genome) has focused on the "classical secretome." This is the collection of proteins that cells actively secrete into the bloodstream, such as hormones or clotting factors. Because these proteins are abundant and easy to detect, they have become the primary targets for drug development. However, this focus creates a blind spot. We have long wondered if expanding our reach to include low-abundance, intracellular proteins could reveal new biological pathways.

The limits of the secretome-centric view

Current high-throughput proteomic platforms, such as the 5,000-protein version of SomaScan, are highly effective at capturing the "loudest" signals in the blood. These are the secreted proteins that are easy to measure reliably. While this has identified many protein quantitative trait loci (pQTLs)—genetic variants that influence protein levels—it leaves a massive portion of the proteome unexplored.

The limitation is not just about quantity. It is about biological coverage. By focusing on the secretome, researchers may miss the subtle regulatory signals of proteins that reside within cells. If a disease is driven by a protein that usually stays inside a cell, current methods might lack the depth to capture its genetic drivers. The authors of this study sought to determine if moving from a 5k to a 7k protein platform improves our ability to perform causal inference (determining if a protein actually causes a disease rather than just being a bystander).

Mapping the expanded genetic architecture

To test this, the researchers used an expanded SomaScan 7k platform. They analyzed 7,144 aptamers (small DNA or RNA molecules used as probes to bind specific proteins) targeting 6,267 proteins. They leveraged data from two large European cohorts, INTERVAL and CHRIS, totaling 13,445 participants. Their methodology moved through several layers of genetic analysis:

  1. pQTL Discovery: They performed genome-wide association studies to identify both cis-pQTLs (variants near the protein-coding gene that regulate it locally) and trans-pQTLs (variants far away that regulate the protein indirectly).
  2. Conditional Analysis: To avoid counting the same genetic signal twice, they used stepwise conditional regression to isolate independent regulatory variants.
  3. Colocalization: They employed Bayesian colocalization to check if the genetic signal for a protein matched the signal for a gene's expression (eQTL) or a disease trait (GWAS). This ensures the association is likely caused by the same underlying biological mechanism.
  4. Mendelian Randomization (MR): Finally, they used these genetic variants as "natural experiments" to estimate the causal effect of protein levels on over 2,000 clinical phenotypes.

As shown in, the expansion fundamentally changed the biological makeup of the dataset.

Figure 1
Figure 1 — from the original paper

The new proteins were heavily enriched for intracellular localization. They were also expected to exist at much lower concentrations in the plasma than the older, secreted proteins.

More proteins, more novel biology

The results of the expansion are significant. The authors report identifying 7,870 significant pQTLs. They used a strict genome-wide significance threshold of $P < 1.26 \times 10^{-11}$ to ensure rigor. Of these, 2,704 (34.3%) had never been reported in five prior large-scale studies. Crucially, over half of these novel associations (53%) involved proteins that were only newly introduced by the 7k platform.

The study also characterizes the complex "traffic jams" of genetic regulation. They identified 22 pleiotropic trans-regulatory hotspots. These are genomic regions that exert broad control over many different proteins simultaneously. These hotspots account for 68% of all trans-pQTLs [Figure 1a]. The researchers found that these hotspots vary greatly. Some, like the GCKR locus, are driven by a single "master regulator" variant [Figure 2a]. Others, like a region on chromosome 3, consist of multiple independent regulatory clusters [Figure 2b].

Regarding disease relevance, the authors used Mendelian randomization to identify 6,340 unique gene–phenotype pairs. This includes many instances where the expansion revealed new disease mechanisms. For example, the study highlights the intracellular protein FN3K, which was newly measured. The researchers found that genetically predicted higher levels of FN3K are associated with lower HbA1c (a marker of long-term blood sugar) and lower type 2 diabetes risk [Figure 3f-g]. This suggests that even "hidden" intracellular proteins can provide actionable clinical insights.

Navigating the noise of low abundance

Despite the gains, the expansion comes with inherent trade-offs. The authors report that newly added proteins were significantly less likely to harbor cis-pQTLs. Specifically, only 15% of newly added proteins had cis-pQTLs, compared to 28% for proteins in the 5k version [Figure 1b]. They argue this is not a failure of assay precision. Instead, it reflects biology. Intracellular proteins are present at much lower concentrations. This makes their genetic signals harder to detect statistically.

There are also technical caveats regarding "epitope-sensitive" associations. Some genetic variants might change the protein's sequence. This can affect how the SomaScan aptamer binds to it. This might change the reading without changing the actual amount of protein in the blood. The authors flag these protein-altering variants (PAVs) as requiring cautious interpretation. They might create "apparent" protein differences that are actually just assay artifacts. Furthermore, the study is limited to populations of European ancestry. This means the specific regulatory hotspots identified here may not apply to all global populations.

The verdict: Expand the search, but watch the signal

Is proteome expansion worth the effort? For researchers looking to move beyond the obvious, the answer is a definitive yes. The study demonstrates that expanding the platform does not just add "more of the same." It adds a non-redundant layer of biology. This allows us to see disease mechanisms that were previously invisible.

However, the transition from the secretome to the intracellular proteome requires a change in strategy. You cannot expect the same high-power, easy-to-find cis-signals you get with abundant, secreted proteins. Instead, practitioners must prepare for lower statistical power. They must also handle a higher prevalence of complex, multi-target regulatory architectures. If you are building a pipeline for drug target discovery, this paper suggests a clear path. While the "low-hanging fruit" remains in the secreted proteins, the next frontier of medicine likely lies in the low-abundance, intracellular signals currently hiding in the noise.

Figures from the paper

Figure 2
Figure 2 — from the original paper
Figure 3
Figure 3 — from the original paper
Figure 4
Figure S1. Number of conditionally independent variants per regional association.
Figure 5
Figure 5 — from the original paper
Figure 6
Figure 6 — from the original paper
Novelty
0.0/10
Overall
0.0/10
#proteomics#GWAS#Mendelian randomization#pQTL#SomaScan
How this was made
Generation

Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: engineering_deepdive
Refinement: 0
Pipeline: forge-1.1

Verification

Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 94% (passed)
Claims verified: 16 / 17

Translation

Model: nvidia/Gemma-4-26B-A4B-NVFP4

Hardware & cost

NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 143,536
Wall-time: 306.5s
Tokens/s: 468.3

Related
Next up

Metabolic Signatures of Healthy Diets Linked to Reduced Risk of MASLD and Cir...

7.6/10· 6 min