Most people still think "genetic disease" means one broken gene. Also, one typo in the code. One villain Small thing, real impact..
That's not how it works for the stuff that actually kills most of us Worth keeping that in mind..
Heart disease. Here's the thing — type 2 diabetes. Alzheimer's. Schizophrenia. Most cancers. Autoimmune conditions. Which means these aren't single-gene disorders. They're not cystic fibrosis or Huntington's or sickle cell. They're something messier — and far more common The details matter here..
What Is a Complex Disease
A complex disease — sometimes called a multifactorial or polygenic disease — is exactly what it sounds like. Practically speaking, multiple genes. Multiple variants. Practically speaking, each one nudging risk up or down by a tiny fraction. Together, they stack up.
But genes aren't the whole story. Environment matters. Which means lifestyle matters. Timing matters. The same genetic hand plays out differently depending on what you eat, how you sleep, what you breathe, whether you smoke, how stressed you are, and a hundred other things.
Here's the thing — there's no clean line between "genetic" and "not genetic." It's all genetic and environmental. The distinction is a false one Worth knowing..
Polygenic vs. Monogenic: The Difference That Matters
Monogenic diseases follow Mendelian rules. One gene, one mutation, predictable inheritance. If you have the mutation, you get the disease (mostly). Plus, think Huntington's. Think Tay-Sachs That's the part that actually makes a difference. Less friction, more output..
Polygenic diseases don't play by those rules. Consider this: each variant might increase risk by 1. 05x or 0.Individually, they're noise. Hundreds — sometimes thousands — of genetic variants contribute. That's why 98x. Collectively, they're signal.
And unlike monogenic conditions, you can't test for a complex disease with a simple yes/no genetic screen. That's not how polygenic risk works.
The "Imperfection" Language Problem
I need to pause on the word "imperfection." It's the user's term. It's common. But it's misleading Simple as that..
Genetic variants aren't "imperfections" in any moral or functional sense. Some increase disease risk in modern environments. Now, they're variations. Some decreased risk in ancestral environments — the thrifty gene hypothesis suggests variants that helped us survive famine now predispose us to diabetes in a world of caloric abundance Simple, but easy to overlook..
Same variant. Different context. Not an imperfection. A mismatch And that's really what it comes down to..
Why It Matters / Why People Care
Because this is what's killing us Took long enough..
Seven of the top ten causes of death in the US are complex diseases. Also, heart disease. But cancer. Stroke. Now, alzheimer's. Diabetes. Chronic lower respiratory disease. Kidney disease And it works..
If you have a family history of any of these, you're not looking at a coin flip. You're looking at a loaded deck — but the loading is subtle, distributed, and modifiable.
The Clinical Reality
Doctors have known this forever. Which means "It runs in the family" is the original polygenic risk score. But medicine has been stuck in a monogenic paradigm because that's what we could test for Not complicated — just consistent..
Now we can genotype hundreds of thousands of variants for $50. Polygenic risk scores (PRS) are entering clinical practice. They're imperfect — more on that — but they're real.
And they change the conversation. A 45-year-old man with a high PRS for coronary artery disease isn't just "at risk.That's why " He's someone who might benefit from a statin now, not in ten years. A woman with high breast cancer PRS might start MRI screening at 30 instead of 40.
This changes depending on context. Keep that in mind.
This is precision prevention. Not precision medicine — prevention. There's a difference.
The Equity Problem Nobody Talks About
Here's what most people miss: almost all GWAS (genome-wide association studies) have been done in European-ancestry populations It's one of those things that adds up..
Polygenic risk scores built on European data don't transfer well to African, Asian, Hispanic, or Indigenous populations. The linkage disequilibrium patterns differ. The allele frequencies differ. The effect sizes differ Less friction, more output..
If we roll out PRS clinically without fixing this, we'll widen health disparities. The people who need precision prevention most will get the least accurate scores.
This isn't a footnote. It's the central challenge of the field right now Small thing, real impact..
How It Works (or How to Do It)
Let's break down the machinery. How do multiple genes actually cause disease?
The Genome-Wide Association Study Pipeline
Step one: recruit tens or hundreds of thousands of people. Genotype them on SNP arrays. Worth adding: cases and controls. Impute the rest But it adds up..
Step two: test each variant for association with the trait. That's why logistic regression for binary traits. Still, linear regression for quantitative ones. Correct for principal components (ancestry), batch effects, relatedness.
Step three: apply a genome-wide significance threshold. Think about it: usually p < 5×10⁻⁸. That's Bonferroni correction for ~1 million independent tests The details matter here..
Step four: replicate. Worth adding: independent cohort. Same direction of effect. Similar magnitude.
Step five: meta-analyze. Combine cohorts. Get better effect estimates Surprisingly effective..
That's the classic GWAS. Not genes. It gives you a list of loci — genomic regions — associated with the trait. On the flip side, not mechanisms. Loci.
From Loci to Biology
This is where it gets hard.
A GWAS hit is usually a tag SNP in linkage disequilibrium with the actual causal variant. An enhancer 500kb away. The causal variant might be in a gene's promoter. But a splice site. In real terms, a non-coding RNA. We often don't know No workaround needed..
Fine-mapping narrows the credible set. Functional genomics — ATAC-seq, Hi-C, eQTL colocalization, CRISPR screens — points to target genes and mechanisms Less friction, more output..
But for most loci? Think about it: we still don't know the gene. Or the mechanism. Consider this: or the cell type. Or the developmental timing.
Polygenic Risk Scores: The Math
Take your GWAS summary statistics. For each variant, you have an effect size (beta) and an allele Not complicated — just consistent..
For an individual, count their effect alleles at each variant. But multiply by the beta. Sum across all variants.
That's a PRS. Simple linear algebra.
But which variants? Practically speaking, that's the "clumping and thresholding" approach — old school. All genome-wide significant ones? Better: LDpred, PRS-CS, SBayesR — Bayesian methods that model linkage disequilibrium and shrink effect sizes Worth knowing..
Even better: incorporate functional annotations. In real terms, variants in enhancers active in relevant tissues. Now, variants in coding regions. Variants with evidence of selection.
The best PRS now explain 10-30% of variance for some traits. Because of that, for coronary artery disease, a top-decile PRS confers ~4x risk compared to bottom decile. That's clinically meaningful Most people skip this — try not to..
But — and this matters — PRS captures common variant heritability. Structural variants. Which means gene-environment interactions. Epigenetics. Plus, rare variants. On the flip side, gene-gene interactions. Most of that isn't in the score yet.
The Missing Heritability Isn't Missing Anymore
Early GWAS found variants explaining tiny fractions of heritability. "Missing heritability" became a meme The details matter here..
Turns out it wasn't missing. Day to day, larger sample sizes found them. It was just distributed across thousands of variants too small to detect individually. The heritability was there all along — just polygenic.
For height, we now explain ~40-50% of variance with common variants. Still, for schizophrenia, ~20-25%. For many diseases, we're still at 10-15%.
The gap that remains? Rare variants. Structural variants. Non-additive effects Simple, but easy to overlook..
And yes, the fraction of heritability that still eludes common‑variant GWAS is increasingly attributed to low‑frequency and rare alleles, structural rearrangements, and context‑dependent effects that escape standard additive models. Whole‑genome sequencing of tens of thousands of individuals has begun to uncover these contributors: rare loss‑of‑function variants in drug‑target genes, copy‑number variations that alter dosage of regulatory landscapes, and non‑additive epistatic networks where the impact of one allele depends on the genetic background of others Easy to understand, harder to ignore..
Statistical frameworks such as burden tests, SKAT-O, and mixed‑model approaches that model both rare and common variants jointly are now standard in sequencing‑based association studies. Functional follow‑up mirrors the GWAS pipeline — CRISPR‑based saturation mutagenesis of non‑coding regions, single‑cell eQTL mapping, and epigenomic profiling in disease‑relevant cell types — to pinpoint whether a rare variant perturbs a promoter, enhancer, splice site, or chromatin architecture.
Equally important is the recognition that gene‑environment interplay can masquerade as missing heritability. On the flip side, interaction‑aware models that incorporate measured exposures (e. g., smoking, diet, pollutants) or proxies such as polygenic scores for environmental sensitivity are beginning to recover additional phenotypic variance, especially for complex traits like asthma, type 2 diabetes, and psychiatric disorders.
It sounds simple, but the gap is usually here.
When rare‑variant burden, structural‑variant maps, epigenetic marks, and interaction terms are integrated with polygenic risk scores, the explained variance for many traits climbs substantially — often approaching the narrow‑sense heritability estimates from twin and family studies. To give you an idea, combined common‑plus‑rare‑variant models for coronary artery disease now explain ~45 % of liability, and for schizophrenia the joint contribution reaches ~35 %.
Conclusion
The journey from a GWAS hit to a mechanistic understanding has evolved from single‑locus tagging to a multilayered, systems‑genetics paradigm. Common variants provide the foundation — polygenic risk scores that capture a substantial proportion of trait variance — while rare variants, structural alterations, epigenetics, and gene‑environment interactions fill the remaining gaps. As sample sizes swell, sequencing technologies become routine, and functional assays scale to single‑cell resolution, the “missing” heritability is less a mystery and more a map waiting to be fully charted. Continued interdisciplinary collaboration — bridging statistics, molecular biology, and clinical phenotyping — will translate these genetic insights into better risk stratification, therapeutic targets, and ultimately, precision medicine.