Posts: 30,062
Threads: 2,135
Joined: Dec 2013
OK, this is going to be very wonky.
The first point is that it is not unusual for the study to be modified while it is running. Sometimes, you can't do things you thought you could do. One example relevant to this study is that originally, patients who had been discharged from the hospital were supposed to come in to be evaluated as outpatients. The problem with this is that it is probably not a good idea to have people who were recently diagnosed with COVID-19 and who are probably still shedding virus to be traveling to a clinic. It puts others at risk of infection, it requires the use of PPE for the clinic staff (which may not be available) and it requires some patients to travel, without knowing whether the patients are able to get to the clinic easily. Once the study staff recognized the problems related to this, they can amend the protocol to say that the final assessment can be done by phone rather than in person.
Changes in the analysis plan are also not unusual. For example, let's say that you were planning to use death as an outcome. Your statistical analysis assumed that 20% of patients in the placebo arm would have died, and the power was based on that assumption. But at the end of your study, only 7% of patients have died, maybe because you recruited healthier patients than you expected. Even in the best of cases (all patients on your study drug survived), you only have 14% of patients on placebo will have died. Now all of a sudden your assumptions about the study are wrong. So maybe you change endpoints, maybe you change the sample size, etc. The FDA allows the final statistical plan be filed after the study starts, as long as it is done before the study is blinded.
So, the original study looks to have used an 8-point ordinal rating as the outcome. The ordinal rating scale captures the patient's care needs on an 8-point scale between 1 (dead) to 8 (home, and back to normal health). The actual scale (from the clinicaltrials.gov site) is as follows:
1) Death; 2) Hospitalized, on invasive mechanical ventilation or extracorporeal membrane oxygenation (ECMO); 3) Hospitalized, on non-invasive ventilation or high flow oxygen devices; 4) Hospitalized, requiring supplemental oxygen; 5) Hospitalized, not requiring supplemental oxygen - requiring ongoing medical care (COVID-19 related or otherwise); 6) Hospitalized, not requiring supplemental oxygen - no longer requires ongoing medical care; 7) Not hospitalized, limitation on activities and/or requiring home oxygen; 8) Not hospitalized, no limitations on activities.
What's great about the scale is that it captures the entire spectrum of disease. A drug could be helpful if more patients are home in the treatment group than in placebo at the primary timepoint (I believe 28 days). It might also be helpful if fewer patients are dead at that timepoint. It might also be that you provide small subtle shifts across the spectrum (a couple fewer patients die, a couple more patients go home, etc) which would give the study more power than using a single endpoint, and it means that you wouldn't be in a situation where your primary endpoint of going home is met, but strangely you also see more patients dying.
There are downsides to this approach, however. The first one is that not all "steps" are created equally. Dead versus critically ill is a much bigger difference than the difference between being in the hospital and needing oxygen versus not needing oxygen. Plus, a bunch of other issues might crop up that are not COVID-related. A patient may be ready to go home but have no place to quarantine and therefore they are kept in the hospital (or conversely, hospitals are so overwhelmed that they send patients home before they would have gone home under ordinary circumstances). Hopefully that gets evened out by randomization, but if it happens a lot it makes your endpoint much less meaningful.
In addition, the statistical analysis of an ordinal endpoint is challenging. The most likely analysis (it wasn't specified) is a proportional odds model. The idea is that you are looking for the proportional odds of being in a higher category on treatment than placebo. The challenge is that the resultant odds ratio is not easily interpretable (odds ratios are challenging in a conventional outcome (yes/no) study, let alone an ordinal endpoint). How do you describe in lay terms what the drug does? Also, if relatively few patients die, there is the issue that if you give patients enough time, most of them (except those who die) will recover. So, if your primary timepoint is pretty far out, you are going to lose a lot of power. At the 15-day mark, which was the old primary endpoint, half the patients on placebo would have been discharged (from reports, median hospitalization is 15 days).
By changing the primary endpoint, the study is easier to digest. Basically, did patients on remdisivir recover (basically defined as being discharged home or being hospitalized but not requiring medical care) faster than patients on placebo? The downside is that it doesn't tell you about the state of people in the hospital. Again, based on reports it sounds like there is "a trend towards" mortality benefit, but the p-value was not statistically significant, and we have no idea of whether it makes a difference in patients in the ICU.
Hope this helps.
BC