Exact accuracy is one metric and it hides the shape of the errors.
Four things fall out of that table. Two are bad for us, one is good, and one is a
claim we made in an earlier draft of this page and have had to withdraw.
Our perfect adjacent accuracy is not a virtue. This is the withdrawn
claim. We wrote 28 of 28 in bold and called it a win. Then we added the third column,
and a classifier that ignores the compound and always answers
“Intermediate” also scores 28 of 28, because a middle answer is never two
classes from anything. Our adjacent accuracy is a consequence of hedging, not of
modelling. It stays in the table because removing it would be worse, but it should not
have been highlighted and it is no longer.
We are badly biased toward the middle. That is the same finding said
plainly. The simulation issues “Intermediate” for 20 of 28 compounds and
manages a single Low call in the entire panel. The hERG margin distributes its calls
far more like the truth. If you want a method that will clear a compound, ours is
currently not it.
Where we do win, it is on under-calls. This is the one column that
goes our way, and it is the column a safety scientist cares about, because a missed
liability costs more than a false alarm. Two compounds are under-called by the
simulation, both High drugs dropped to Intermediate, and neither reaches Low. The hERG
margin under-calls four, and two of those go all the way to Low:
astemizole and risperidone. Astemizole was withdrawn from the market
over QT prolongation, its free Cmax is 0.26 nM against a hERG IC50
of 33.3 nM, and a 128-fold margin reads as safe on a spreadsheet.
The full read on astemizole is here.
Verapamil is not the win we first described. It is the one drug the
hERG margin misses by two classes: clinically Low, called High, because verapamil
blocks calcium hard enough to offset its own repolarization delay and a hERG-only score
cannot see the offset. But we do not get verapamil right either. We call it
Intermediate. The margin fails it by 0.038 log units against a cut point fitted on
twelve drugs, and a shift of 0.04 in that cut (about seven percent) erases the error
entirely. Head to head, the simulation is right on two compounds the
margin misses (astemizole, risperidone) and wrong on three the margin gets
(loratadine, nifedipine, nitrendipine). Net, we are one behind. Verapamil is a good
illustration of what multichannel simulation is for. It is not evidence that
ours works.
Neither of us is the published state of the art. The ORd/qNet column is
the CiPA consortium’s own result and its High-vs-rest separation is perfect. Note
the label: those two figures are reported on CiPA’s own validation subset, not on
all 28, so the column is not on the same footing as ours. The published 1.00 carries a
95% confidence interval of 0.92 to 1. We reproduce the column here rather than omit
it.