Following on from the discussion of sampling variability: if simultaneous biopsies from the same liver disagree by a fibrosis stage in a meaningful fraction of cases, and if ballooning has only moderate inter-observer agreement, then the measurement error on the primary endpoint of every MASH trial is large. I would expect that to make these trials hopeless.
Yet they do detect effects, sometimes with substantial separations - resolution rates in the sixties against placebo in the thirties. I would like to understand how that is possible. Is the noise smaller than I am imagining, does randomisation handle it in some way I am not seeing, or are the effect sizes simply big enough to survive it?
And a design question: given all of this, what would you actually change about how these trials are run? I have seen suggestions about central reading, artificial-intelligence-assisted scoring, and continuous rather than categorical endpoints, and I cannot tell which of those addresses the real problem.