Most people still think "genetic disease" means one broken gene. But one typo in the code. One villain Not complicated — just consistent..
That's not how it works for the stuff that actually kills most of us.
Heart disease. Type 2 diabetes. Now, alzheimer's. So schizophrenia. Most cancers. Also, autoimmune conditions. These aren't single-gene disorders. They're not cystic fibrosis or Huntington's or sickle cell. They're something messier — and far more common Easy to understand, harder to ignore..
What Is a Complex Disease
A complex disease — sometimes called a multifactorial or polygenic disease — is exactly what it sounds like. On top of that, multiple genes. Worth adding: multiple variants. That said, each one nudging risk up or down by a tiny fraction. Together, they stack up.
But genes aren't the whole story. Environment matters. Day to day, lifestyle matters. Day to day, timing matters. The same genetic hand plays out differently depending on what you eat, how you sleep, what you breathe, whether you smoke, how stressed you are, and a hundred other things Easy to understand, harder to ignore..
Here's the thing — there's no clean line between "genetic" and "not genetic." It's all genetic and environmental. The distinction is a false one.
Polygenic vs. Monogenic: The Difference That Matters
Monogenic diseases follow Mendelian rules. On top of that, if you have the mutation, you get the disease (mostly). Even so, think Huntington's. Here's the thing — one gene, one mutation, predictable inheritance. Think Tay-Sachs.
Polygenic diseases don't play by those rules. 98x. Each variant might increase risk by 1.Individually, they're noise. But 05x or 0. Hundreds — sometimes thousands — of genetic variants contribute. Collectively, they're signal.
And unlike monogenic conditions, you can't test for a complex disease with a simple yes/no genetic screen. That's not how polygenic risk works Simple, but easy to overlook..
The "Imperfection" Language Problem
I need to pause on the word "imperfection." It's the user's term. It's common. But it's misleading.
Genetic variants aren't "imperfections" in any moral or functional sense. On top of that, they're variations. Some increase disease risk in modern environments. Some decreased risk in ancestral environments — the thrifty gene hypothesis suggests variants that helped us survive famine now predispose us to diabetes in a world of caloric abundance Small thing, real impact..
Same variant. Different context. Not an imperfection. A mismatch.
Why It Matters / Why People Care
Because this is what's killing us That's the part that actually makes a difference. That's the whole idea..
Seven of the top ten causes of death in the US are complex diseases. In real terms, heart disease. Cancer. Stroke. On the flip side, alzheimer's. Diabetes. Consider this: chronic lower respiratory disease. Kidney disease Nothing fancy..
If you have a family history of any of these, you're not looking at a coin flip. You're looking at a loaded deck — but the loading is subtle, distributed, and modifiable.
The Clinical Reality
Doctors have known this forever. Consider this: "It runs in the family" is the original polygenic risk score. But medicine has been stuck in a monogenic paradigm because that's what we could test for Surprisingly effective..
Now we can genotype hundreds of thousands of variants for $50. That's why polygenic risk scores (PRS) are entering clinical practice. They're imperfect — more on that — but they're real.
And they change the conversation. On top of that, a 45-year-old man with a high PRS for coronary artery disease isn't just "at risk. " He's someone who might benefit from a statin now, not in ten years. A woman with high breast cancer PRS might start MRI screening at 30 instead of 40 Practical, not theoretical..
This is precision prevention. In real terms, not precision medicine — prevention. There's a difference Worth keeping that in mind..
The Equity Problem Nobody Talks About
Here's what most people miss: almost all GWAS (genome-wide association studies) have been done in European-ancestry populations And that's really what it comes down to..
Polygenic risk scores built on European data don't transfer well to African, Asian, Hispanic, or Indigenous populations. But the linkage disequilibrium patterns differ. The allele frequencies differ. The effect sizes differ Turns out it matters..
If we roll out PRS clinically without fixing this, we'll widen health disparities. The people who need precision prevention most will get the least accurate scores That alone is useful..
This isn't a footnote. It's the central challenge of the field right now.
How It Works (or How to Do It)
Let's break down the machinery. How do multiple genes actually cause disease?
The Genome-Wide Association Study Pipeline
Step one: recruit tens or hundreds of thousands of people. Still, cases and controls. Genotype them on SNP arrays. Impute the rest.
Step two: test each variant for association with the trait. Logistic regression for binary traits. Because of that, linear regression for quantitative ones. Correct for principal components (ancestry), batch effects, relatedness.
Step three: apply a genome-wide significance threshold. Plus, usually p < 5×10⁻⁸. That's Bonferroni correction for ~1 million independent tests.
Step four: replicate. Independent cohort. Same direction of effect. Similar magnitude That alone is useful..
Step five: meta-analyze. Combine cohorts. Get better effect estimates The details matter here..
That's the classic GWAS. It gives you a list of loci — genomic regions — associated with the trait. Not genes. Not mechanisms. Loci Took long enough..
From Loci to Biology
This is where it gets hard.
A GWAS hit is usually a tag SNP in linkage disequilibrium with the actual causal variant. The causal variant might be in a gene's promoter. An enhancer 500kb away. A splice site. A non-coding RNA. We often don't know.
Fine-mapping narrows the credible set. Functional genomics — ATAC-seq, Hi-C, eQTL colocalization, CRISPR screens — points to target genes and mechanisms Worth knowing..
But for most loci? Or the mechanism. We still don't know the gene. Or the cell type. Or the developmental timing.
Polygenic Risk Scores: The Math
Take your GWAS summary statistics. For each variant, you have an effect size (beta) and an allele Less friction, more output..
For an individual, count their effect alleles at each variant. Multiply by the beta. Sum across all variants.
That's a PRS. Simple linear algebra That's the part that actually makes a difference..
But which variants? Also, all genome-wide significant ones? Plus, that's the "clumping and thresholding" approach — old school. Better: LDpred, PRS-CS, SBayesR — Bayesian methods that model linkage disequilibrium and shrink effect sizes.
Even better: incorporate functional annotations. Still, variants in coding regions. Variants in enhancers active in relevant tissues. Variants with evidence of selection.
The best PRS now explain 10-30% of variance for some traits. For coronary artery disease, a top-decile PRS confers ~4x risk compared to bottom decile. That's clinically meaningful Simple, but easy to overlook..
But — and this matters — PRS captures common variant heritability. Consider this: epigenetics. Gene-gene interactions. Even so, structural variants. In practice, rare variants. Day to day, gene-environment interactions. Most of that isn't in the score yet.
The Missing Heritability Isn't Missing Anymore
Early GWAS found variants explaining tiny fractions of heritability. "Missing heritability" became a meme The details matter here..
Turns out it wasn't missing. It was just distributed across thousands of variants too small to detect individually. Larger sample sizes found them. The heritability was there all along — just polygenic Still holds up..
For height, we now explain ~40-50% of variance with common variants. For schizophrenia, ~20-25%. For many diseases, we're still at 10-15%.
The gap that remains? Rare variants. Structural variants. Non-additive effects.
And yes, the fraction of heritability that still eludes common‑variant GWAS is increasingly attributed to low‑frequency and rare alleles, structural rearrangements, and context‑dependent effects that escape standard additive models. Whole‑genome sequencing of tens of thousands of individuals has begun to uncover these contributors: rare loss‑of‑function variants in drug‑target genes, copy‑number variations that alter dosage of regulatory landscapes, and non‑additive epistatic networks where the impact of one allele depends on the genetic background of others But it adds up..
Statistical frameworks such as burden tests, SKAT-O, and mixed‑model approaches that model both rare and common variants jointly are now standard in sequencing‑based association studies. Functional follow‑up mirrors the GWAS pipeline — CRISPR‑based saturation mutagenesis of non‑coding regions, single‑cell eQTL mapping, and epigenomic profiling in disease‑relevant cell types — to pinpoint whether a rare variant perturbs a promoter, enhancer, splice site, or chromatin architecture Worth keeping that in mind..
Equally important is the recognition that gene‑environment interplay can masquerade as missing heritability. And interaction‑aware models that incorporate measured exposures (e. g., smoking, diet, pollutants) or proxies such as polygenic scores for environmental sensitivity are beginning to recover additional phenotypic variance, especially for complex traits like asthma, type 2 diabetes, and psychiatric disorders Simple as that..
When rare‑variant burden, structural‑variant maps, epigenetic marks, and interaction terms are integrated with polygenic risk scores, the explained variance for many traits climbs substantially — often approaching the narrow‑sense heritability estimates from twin and family studies. To give you an idea, combined common‑plus‑rare‑variant models for coronary artery disease now explain ~45 % of liability, and for schizophrenia the joint contribution reaches ~35 % That alone is useful..
Conclusion
The journey from a GWAS hit to a mechanistic understanding has evolved from single‑locus tagging to a multilayered, systems‑genetics paradigm. Common variants provide the foundation — polygenic risk scores that capture a substantial proportion of trait variance — while rare variants, structural alterations, epigenetics, and gene‑environment interactions fill the remaining gaps. As sample sizes swell, sequencing technologies become routine, and functional assays scale to single‑cell resolution, the “missing” heritability is less a mystery and more a map waiting to be fully charted. Continued interdisciplinary collaboration — bridging statistics, molecular biology, and clinical phenotyping — will translate these genetic insights into better risk stratification, therapeutic targets, and ultimately, precision medicine Most people skip this — try not to..