Status of this manuscript
This is an unpublished draft. It has not been peer reviewed, IJRM-SSS has not published its first issue, and nothing here should be cited as a published finding. It is posted so that the argument can be criticised before it is fixed in print. Authorship, competing interests and the handling arrangement described in the Institute’s conflict-of-interest disclosure apply.
Abstract
The quality-adjusted life year multiplies a health-state utility by time. That single design choice makes it a population instrument of enormous utility and an individual instrument of poor validity: a benefit that returns quality without extending life is scaled by a small number and disappears. Survivorship and surveillance medicine sits almost entirely in that blind spot. We review the QALY’s construction, the discrimination critique that culminated in a statutory prohibition on its use in US Medicare drug-price negotiation, and the three principal alternatives advanced in response — equal value of life-years gained (evLYG), health years in total (HYT), and generalized risk-adjusted cost-effectiveness (GRACE) — including published demonstrations that two of the three can produce logically inconsistent rankings. We then set out the iQALY, an individual-level metric in which quality returned is the achievement and duration is a co-equal term rather than a multiplier, and state the conditions under which it should be rejected.
1. The problem this paper is about
A woman finishes treatment for a lymphoma at 34 and is cured. Twenty years of anthracycline-exposed myocardium lie ahead of her. Serial echocardiography with strain imaging can identify declining function while it is still reversible. It will not, for most people in her position, add a single year of life — it will decide whether those years are spent in heart failure.
Under a quality-adjusted life year, the value of that surveillance is a quality decrement avoided, multiplied by time, discounted, and spread across every person screened to find the few who benefit. It approaches zero. The metric is not malfunctioning; it is doing exactly what it was built to do. The question this paper asks is whether it is the right instrument for the decision, and what a companion instrument would have to look like.
2. Where the QALY came from, and what it was for
The QALY was constructed to solve a real and hard problem: how to compare interventions across unlike diseases with a single unit, so that a fixed budget could be allocated to produce the most health. Its form — utility weight × time — makes that comparison possible. Generic preference-based instruments such as the EQ-5D supply the weights; national value sets make them comparable; a cost-per-QALY threshold turns them into a decision.
It works. Health technology assessment agencies across Europe, the UK, Canada and Australia use it, and the discipline of forcing an explicit, auditable trade-off has almost certainly improved allocation over the alternative of unexamined judgement. Any argument for a companion metric has to begin by conceding this. The QALY is not a bad instrument; it is an instrument with a domain.
Its comparability across diseases has also been examined against the disability-adjusted life year, with empirical work finding that the choice between them can shift results enough to change a decision at conventional thresholds, though not systematically in one direction1.
3. The discrimination critique, and the US legal response
The multiplicative form has a consequence its designers did not intend. If a person already lives with reduced health-related quality of life — from disability, chronic illness, or advanced age — then any life-year saved for that person is worth less in the numerator than the same year saved for someone healthier. The metric does not merely measure the disadvantage; it propagates it into the allocation.
In the United States this moved from academic objection to statute. The Inflation Reduction Act prohibits the Centers for Medicare and Medicaid Services from using standard quality-adjusted life-years, or other value-assessment methods that discriminate against the aged, terminally ill, or disabled, when setting maximum fair prices for prescription drugs2. Whatever one thinks of that policy, its practical effect is that the largest payer in the United States needs a value framework that is not the QALY — and the search for one is now an active field rather than a thought experiment3.
4. Prior attempts at the same problem
| Method | Core idea | Published objection |
|---|---|---|
| evLYG equal value of life-years gained | Value each life-year gained equally regardless of the recipient’s health state, removing the penalty for pre-existing disability. In use at ICER. | Fails to credit quality-of-life gains during added years; can produce an unstable ranking of options4. |
| HYT health years in total | Separate life-expectancy change and quality change onto an additive scale rather than multiplying them; same axiomatic foundations as the QALY5. | Can violate independence of irrelevant alternatives; assumes separability of quality and life-years; requires counterfactual quality of life for the dead4. |
| GRACE generalized risk-adjusted CEA | Derive severity and disability adjustments from expected-utility microeconomic foundations rather than ad hoc ethical rules6. | Proposed as the principled alternative to evLYG and HYT2; demanding to parameterise, and not yet embedded in routine coverage practice. |
| Shortfall weighting absolute, proportional, fair innings | Adjust the willingness-to-pay threshold by how much health the illness takes away. | The three disagree with one another materially, so at most one can describe patient preferences; stair-step brackets raise their own ethical problems6. |
5. The lesson we take from those attempts
This history is the strongest argument against the metric we are about to propose, and we would rather make it ourselves. Two of the three leading alternatives were advanced by serious health economists to fix precisely the problem we are describing, and both were subsequently shown to generate logically inconsistent decisions — rankings that flip when an irrelevant option is added, or that value a survival gain negatively4. Good intentions about non-discrimination did not protect them.
The obligation this places on the iQALY is specific: it must be tested for axiomatic consistency, not merely for whether its answers feel more humane. A metric that produces the intuitively right answer by an incoherent route will be discarded, and should be.
6. What the iQALY proposes
The iQALY is not a replacement for the QALY and does not attempt population allocation. It is an individual metric for a different question: what does this intervention return to this person?
Its structural departure is to stop treating duration as a multiplier. Quality returned (Δq) is the achievement. Duration enters as an endurance term (H) alongside a reference horizon (R × Href), additively rather than multiplicatively, so that quality returned over a short horizon scores at parity with quality returned over a long one. A profound benefit delivered in the last nine months of a life is not discounted for arriving late. There is no duration amplifier and no discount rate.
The consequences are deliberate. A prevented decrement counts as much as a gain. A patient with a short prognosis is not thereby worth less. And an intervention whose entire value is quality — the six surveillance tests catalogued on the Institute’s iQALY page — can clear a threshold it currently cannot approach.
7. Coverage pathways: the state of the art, and where an iQALY would enter
Coverage decisions today are made at population level and appealed at individual level. The appeal is the only routine point in the system where an individual argues against an average — and it is conducted almost entirely without a quantitative instrument, on narrative and clinical letters.
That asymmetry is the opening. We propose the iQALY enter through appeal rather than through allocation: a declared Δq, H and R with assumptions on the page, replacing a letter of hardship with a structured claim that can be audited and, crucially, checked against what actually happened. If appeals reliably identify the same kinds of patients, and follow-up shows those patients realised the quality argued for, the metric earns standing as a prospective justification. Only then does the question of a standard for overruling a population determination become answerable.
None of the alternatives reviewed above takes this route; all are proposed as replacements at the allocation step, which is where they meet the most resistance and where their inconsistencies bite hardest.
8. What it demands: individual prognosis
An individual metric is only as good as individual prediction. H is a prognosis, and population cost-effectiveness can average past a poor prognostic estimate in a way that an individual decision cannot. The iQALY therefore raises, rather than lowers, the burden on prognostication, and any implementation must carry its uncertainty explicitly rather than reporting a point estimate.
A survivorship clinic with longitudinal follow-up, recorded toxicity and recurrence data is the natural instrument for estimating the two quantities the metric needs in an individual: the probability of a specific quality decrement, and the probability of retaining the time the population curve implies. Building that estimation is the empirical programme; it has not been done, and this paper claims no results from any cohort.
9. Limitations, and how this proposal could fail
- It may be axiomatically inconsistent. It has not yet been subjected to the analysis that falsified HYT and evLYG. That analysis should be done by someone with no stake in the answer.
- Δq may not be measurable at individual level with the instruments available. Generic utility instruments are insensitive to several of the decrements this metric is meant to capture.
- It could be gamed. Any appeal instrument with declared parameters invites optimistic declaration. Without audit against realised outcomes it becomes advocacy with arithmetic.
- It does not solve allocation. If every individual appeal succeeds, the budget constraint reappears elsewhere. We do not have an answer to this and do not claim one.
- Conflict of interest. The authors are affiliated with a practice that delivers the kind of surveillance this metric would favour. That is disclosed, and it is a reason for external replication rather than internal validation.
References
- Augustovski F, Colantonio LD, Galante J, et al. Measuring the benefits of healthcare: DALYs and QALYs — does the choice of measure matter? Int J Health Policy Manag. 2018;7(2):120–136. doi:10.15171/ijhpm.2017.47
- Lakdawalla DN, Doctor JN. A principled approach to non-discrimination in cost-effectiveness. Eur J Health Econ. 2024;25(8):1393–1416. doi:10.1007/s10198-023-01659-7
- DiStefano MJ, Zemplenyi A, Anderson KE, et al. Alternative approaches to measuring value: an update on innovative methods in the context of the United States Medicare drug price negotiation program. Expert Rev Pharmacoecon Outcomes Res. 2024;24(2):171–180. doi:10.1080/14737167.2023.2283584
- Paulden M, Sampson C, O’Mahony JF, et al. Logical inconsistencies in the health years in total and equal value of life-years gained. Value Health. 2024;27(3):356–366. doi:10.1016/j.jval.2023.11.009
- Basu A, Carlson J, Veenstra D. Health years in total: a new health objective function for cost-effectiveness analysis. Value Health. 2020;23(1):96–103. doi:10.1016/j.jval.2019.10.014
- Phelps CE, Lakdawalla DN. Methods to adjust willingness-to-pay measures for severity of illness. Value Health. 2023;26(7):1003–1010. doi:10.1016/j.jval.2023.02.001
