finding · the benchmark

We publish the benchmark we lose

On the 28-drug CiPA panel, a hERG margin built from two published numbers classifies torsade risk more accurately than our simulation does: 18 of 28 against 17 of 28. The comparison is on our site, with every number behind it.

The loss is deliberate. A tool that sells a mechanism read should be able to show it beats the cheap alternative, and the comparison, win or loss, stays on the site. We froze the thresholds before scoring the held-out drugs and put the results up.

It is worse than a one-drug loss. The reference file carries two published hERG IC50 columns that disagree by up to seventeen-fold on a single drug. We use the static column everywhere. Feed the margin the dynamic column instead and it scores 19 of 28, two ahead of us.

three-class accuracy, CiPA-28
iPSC-CM simulationhERG-only marginalways "Intermediate"
12 training drugs8 of 128 of 124 of 12
16 blind drugs9 of 16 (0.56)10 of 16 (0.63)7 of 16 (0.44)
all 2817 of 28 (0.61)18 of 28 (0.64)11 of 28 (0.39)

Two published numbers and a division beat a CVODE solve per concentration per compound. On this task, on this panel, by this metric. The third column matters too: a classifier that always answers "Intermediate" gets 11 of 28. Both real methods clear it, but not by the margin either of us would like.

Where each method is wrong

Exact accuracy hides the shape of the errors. Head to head, our call is correct on two drugs the margin misses (astemizole and risperidone) and wrong on three it gets (loratadine, nifedipine, nitrendipine). Net, we are one behind. The margin makes a two-class error on verapamil, the one compound where any method on the panel misses by that much.

The margin wins on the aggregate and loses on the cases a cardiac-safety team cares most about. Astemizole, a drug withdrawn for TdP, is read as Low by the margin. Verapamil, a safe calcium blocker, is read as High. Astemizole is the wedge argument. Verapamil is the multichannel one. The mechanism read gets both classes closer to the truth even where it lands one class off.

The gap between 17 and 18 of 28 is one compound, a snapshot against one published panel, scored by one metric. It is published so that nobody has to take our word for where the method stands.

limits for a buyer
the benchmark, in full

The complete comparison, including the three numbers that make the loss look worse than it first read, is on the benchmark page. Read every number.

see what we do instead more writing
CiPA 28-drug reference set; thresholds fitted on 12 training drugs and frozen before scoring the 16 held-out. Harness uses the static Li et al. 2017 IC50 column. Simulation: Kernik-Clancy 2019 human iPSC-CM model (doi.org/10.1113/JP277724) integrated with Myokit/CVODE. Full method on the benchmark page.
get in touch
[email protected]