Blood Test Flags 14 Proteins to Predict Lung-cancer Risk up to Five Years Early
Why in News?
Two strands of cancer research reported within days of each other in June 2026 point to the same promise — a simple blood draw that flags who is likely to develop lung cancer long before any tumour shows on a scan. The headline finding, in the journal Cell, identifies a signature of 14 plasma proteins that can predict a lung-cancer diagnosis up to five years before it happens.
- The Cell study (Francis Crick Institute and UCL, Charlie Swanton’s group, 80+ collaborators) used machine learning on 48,000+ UK Biobank participants.
- The 14-protein signature was validated across eight global datasets, including a cohort of people who never smoked.
- The signature reflects an inflamed, pre-cancer lung environment linked to air-pollution-driven interleukin-1 beta (IL-1 beta) signalling — not an existing tumour.
- A separate Cleveland Clinic and DELFI Diagnostics blood test reads cell-free DNA (cfDNA) fragment patterns — a technique called fragmentomics.
- The DELFI-L101 study reported an ROC-curve value of about 0.81 across 294 lung-cancer and 661 non-cancer participants.
- Both predict risk or aid early detection; neither is a confirmed standalone diagnostic test.
The development matters in the context of:
- India’s high lung-cancer burden, a large share of it in non-smokers, with chronic air pollution a major suspected driver.
- Current standard screening — low-dose CT (LDCT) — aimed mainly at older heavy smokers, leaving many at-risk groups uncovered and badly under-used.
UPSC Relevance
Prelims Relevance
- Biomarker — any measurable biological signal (protein, gene, DNA fragment) that tracks with a disease or future risk.
- Liquid biopsy — looking for disease signals in a blood sample rather than removing tissue.
- Cell-free DNA (cfDNA) — short DNA shed into blood by dying cells; the tumour-derived portion is circulating tumour DNA (ctDNA).
- Fragmentomics (the DELFI approach) — reads the pattern in which cfDNA is broken up, since cancer alters DNA fragmentation.
- 14-protein signature (Cell study) — predicts lung-cancer risk up to five years before diagnosis.
- Machine learning on 48,000+ UK Biobank participants; validated across eight global datasets.
- Interleukin-1 beta (IL-1 beta) — an inflammation messenger; the signature reflects pollution-linked inflamed lung environment.
- Sensitivity — share of true cases a test correctly flags; specificity — share of healthy people it correctly clears.
- Lead time — how much earlier than usual a test detects disease before symptoms or imaging.
- Low-dose computed tomography (LDCT) — current standard lung-cancer screening, aimed mainly at older heavy smokers.
- DELFI-L101 study — ROC-curve value of about 0.81 (1.0 perfect, 0.5 no better than chance).
Mains Relevance
GS Paper 3 (Science and Technology):
- A clean illustration of biomarkers, liquid biopsy, cfDNA, ctDNA, fragmentomics and proteomics, and how machine learning is reshaping early disease detection.
- The shift from finding cancer to forecasting it — detection toward prediction and potentially prevention.
GS Paper 3 / GS Paper 2 (Public health and pollution):
- India’s high lung-cancer burden, much of it in non-smokers, and the inflammatory damage chronic air pollution does to the lung.
- Reframing air pollution as a medical, not just an environmental, problem — pollution control as a form of cancer prevention.
Essay
- Equity and access — turning a laboratory signature into an affordable, validated, made-in-India screening tool on a system that struggles with basic cancer screening.
Background and Context
Three durable anchors — biomarkers, liquid biopsy, early detection
- A biomarker is any measurable biological signal whose level or pattern tracks with a disease or future risk — the 14-protein signature is a protein biomarker; the DELFI test reads a DNA-based one.
- Liquid biopsy is the broader strategy of looking for these signals in an ordinary blood sample rather than cutting out tissue.
- A traditional biopsy removes tissue; a liquid biopsy reads what the tumour or at-risk tissue sheds into the blood — cheaper, repeatable, far less invasive, ideal for screening large populations.
The DNA side — cfDNA and fragmentomics
- Cells throughout the body die and release short stretches of DNA into the blood; the tumour-derived portion is circulating tumour DNA (ctDNA).
- Early cfDNA tests tried to read actual cancer mutations — hard when a tumour is tiny and sheds little.
- Fragmentomics sidesteps that: it studies how cfDNA is chopped up (the lengths and positions of fragments), because cancer disturbs how DNA is packaged in cells.
- A machine-learning model learns to recognise the cancer-associated pattern across the whole genome, using cheap low-coverage sequencing rather than expensive deep reads.
What early detection means, and how it is judged
- Sensitivity — the share of people who truly have/will develop the disease that the test correctly flags; low sensitivity misses real cases.
- Specificity — the share of healthy people the test correctly clears; low specificity floods the system with false alarms.
- Lead time — the head start a test buys; a five-year lead time is striking because lung cancer is so often found late.
- LDCT, the current standard, works but is narrow by design (older heavy smokers) and badly under-used — only a small minority of eligible people complete the scan, fewer than 40% return for the annual repeat.
- A blood-based risk test would not replace the CT — it would decide, more cheaply and widely, who should be sent for one, reaching non-smokers and the pollution-exposed.
The Cell protein signature — what it found
- Machine learning on blood-plasma protein data from 48,000+ UK Biobank participants converged on a signature of 14 proteins flagging a future diagnosis up to five years in advance.
- The signature held even in a never-smoker cohort — the group conventional screening tends to miss.
- The proteins reflect an altered, inflamed lung environment that precedes cancer; air pollution can drive this inflammation, partly by raising IL-1 beta signalling.
- The same signature also appeared in people who later developed idiopathic pulmonary fibrosis and COPD — suggesting it reads a broader state of lung injury.
- In a re-analysis of the CANTOS trial (4,651 participants), people with a high baseline signature gained most from a drug blocking IL-1 beta — their lung-cancer risk almost halved — hinting the biology could be a prevention target.
The DELFI fragmentomics test — what it found
- Cleveland Clinic with DELFI Diagnostics reported a blood test that measures no proteins — it reads cfDNA fragment patterns.
- Cancer makes cfDNA more varied in size and pattern; machine learning is trained to tell sick from healthy.
- DELFI-L101: 294 people with lung cancer, 661 without; ROC-curve value about 0.81, called accurate enough to pursue toward practice.
- Its real aim is reach — pulling more people into the screening funnel via an easy blood draw, not higher accuracy than the CT itself.
- By design it recruited adults aged 50+ with a heavy smoking history of at least 20 pack-years — the group already eligible for CT screening.
Two tools, two different questions
- Protein signature — built for risk prediction in apparently healthy people: who is likely to develop lung cancer over the next few years; deliberately reached beyond smokers.
- cfDNA fragmentomics — built mainly as a screening triage: who among eligible adults should be sent for an imaging scan now.
- Neither is a diagnosis — a positive result still leads to a low-dose CT and tissue confirmation.
The India fit — and the caveat
- The dominant worldwide screening criterion — heavy smoking plus older age — was built for Western, smoking-driven lung cancer.
- India’s pattern differs: a substantial share of cases occur in people who never smoked, with chronic outdoor and household air pollution a major suspected driver — the very inflammatory pathway the 14-protein signature reads.
- A risk-prediction test that works in non-smokers and reflects pollution-linked inflammation is, in principle, far better matched to India than a smoking-history rule.
- The structural caveat is access — such tests arrive expensive and patent-protected; turning a lab signature into an affordable, validated, made-in-India tool is the harder, longer task. Links to broader preventive-health themes in our note on a decade of the PMSMA maternal-health programme.
Challenges and concerns
- Research stage, not approved — both tools need large prospective trials (DELFI talks of next-stage validation in hundreds, then 15,000 participants).
- False-positive burden — in a population where lung cancer is still relatively uncommon, imperfect specificity flags many healthy people, driving needless scans and anxiety; an ROC of ~0.81 is a starting point, not proof of clinical readiness.
- Cohort bias — a model built mainly on UK Biobank and Western smoker cohorts may not transfer to India’s largely non-smoking, pollution-exposed cases without local validation.
- Cost and access — advanced proteomics and genome-sequencing assays are expensive and patent-protected.
- Confirmation still needed — a positive blood result triggers imaging and tissue biopsy, not a final diagnosis, so the downstream care pathway must exist.
Way Forward
Validate before roll-out
- Move the protein signature and fragmentomics tests through large prospective trials, including Indian cohorts, so accuracy claims are validated in real screening conditions rather than retrospective data.
Build the downstream pathway
- Affordable low-dose CT capacity and confirmatory biopsy, so flagging more at-risk people translates into earlier treatment rather than dead-end positives.
Invest in indigenous capacity and prevention
- Lower-cost domestic proteomics and sequencing capacity, plus pollution control as primary prevention — tackling both the cause of pollution-linked lung cancer and the means to detect it early.
Conclusion
The real shift is from finding cancer to forecasting it — catching the inflamed, at-risk lung before a tumour forms, where treatment is cheaper and survival far higher. The CANTOS re-analysis hints the same biology could be acted on, not just observed.
The caveats are as exam-worthy as the promise: these are research findings, validated retrospectively, that must prove themselves prospectively, and any screening tool runs into the arithmetic of sensitivity and specificity.
For India the pointed lesson is fit — a test that works in non-smokers and reads pollution-linked inflammation suits the disease pattern, but affordability is the task that policy, not the lab, must solve.
UPSC Practice Questions
Prelims MCQ 1
With reference to liquid biopsy and biomarker-based testing, consider the following statements:
- Cell-free DNA (cfDNA) is short DNA shed into the blood by dying cells.
- The tumour-derived portion of cfDNA is called circulating tumour DNA (ctDNA).
- Fragmentomics reads the pattern in which cfDNA is broken up, rather than the genetic code letter by letter.
How many of the above statements are correct?
(a) Only one (b) Only two (c) All three (d) None
Answer: (c)
Explanation:
- cfDNA is short DNA shed into the blood by dying cells — correct.
- The tumour-derived fraction is ctDNA — correct.
- Fragmentomics studies cfDNA fragmentation patterns rather than reading mutations directly — correct.
Prelims MCQ 2
In disease screening, “specificity” refers to which of the following?
(a) The share of people with the disease that the test correctly flags (b) The share of healthy people the test correctly clears (c) How much earlier than usual a test detects disease (d) The proportion of false-positive results among all positives
Answer: (b)
Specificity is the share of healthy people a test correctly clears; sensitivity (option a) is the share of true cases it flags; lead time (option c) is the head start a test buys. Low specificity drives false alarms and needless scans.
UPSC Mains Questions
Liquid biopsy and biomarker-based blood tests are shifting cancer care from detection toward prediction and prevention. Discuss the underlying science and the challenges of translating such tools into population screening. (GS3, 15 marks, 250 words)
A large share of India’s lung-cancer cases occur in non-smokers, with air pollution a major suspected driver. Examine why smoking-history-based screening is poorly suited to India and how new biomarker tests might help. (GS3, 15 marks, 250 words)
What did the new lung-cancer study find?
A study in the journal Cell, led by the Francis Crick Institute and UCL, identified 14 proteins in blood plasma whose levels can flag a future lung-cancer diagnosis up to five years before it occurs. The signature was found using machine learning on over 48,000 UK Biobank samples and held up across eight datasets worldwide. It predicts risk, not a confirmed diagnosis.
What is a liquid biopsy?
A liquid biopsy looks for signs of cancer in an ordinary blood sample instead of cutting out tissue. It reads biomarkers, proteins, or fragments of DNA, that a tumour or an at-risk tissue sheds into the bloodstream. Because it is cheap, repeatable and minimally invasive, it is well suited to screening large populations who would never undergo a tissue biopsy.
What is fragmentomics and how does DELFI use it?
Fragmentomics studies how cell-free DNA, the loose DNA floating in blood, is broken up. Cancer disturbs the way DNA is packaged in cells, so it changes the size and pattern of these fragments. The DELFI test, developed with Cleveland Clinic, uses machine learning to spot the cancer-associated fragmentation pattern across the genome, aiming to widen who gets screened.
Is this a confirmed lung-cancer diagnostic test?
No. Both the 14-protein signature and the DELFI fragmentomics test are research findings that predict risk or flag who needs further checks. A positive result still leads to a low-dose CT scan and a tissue biopsy to confirm cancer. The tests need large prospective trials before they can be used routinely in clinical care.
Why does this matter for India?
Lung cancer is among India’s leading cancers, and a large share of cases occur in people who never smoked, with air pollution a major suspected cause. Standard screening targets older heavy smokers and so misses many Indian patients. A blood test that works in non-smokers and reads pollution-linked lung inflammation could fit India’s disease pattern far better, if it can be made affordable.
How is this different from a low-dose CT scan?
A low-dose CT (LDCT) scan is the current standard screening, an imaging test aimed mainly at older heavy smokers, but it is expensive and badly under-used. The new blood tests do not replace it; they aim to decide, more cheaply and across wider groups, who should be sent for a CT, and could catch risk years before a tumour shows on any scan.