When a free hERG margin beats us.

On the public 28-drug CiPA panel, a hERG-only safety margin built from two published numbers classifies torsade risk more accurately than our simulation does. 18 of 28 against our 17 of 28, and 19 of 28 if you feed it the other published IC50 column. We ran the comparison, we lost it, and this page publishes the numbers behind every claim on it, including the three that make the loss look worse than we first wrote it up.

The hERG margin beats our simulation

The comparison is a three-class call, Low, Intermediate or High torsade risk, on the 28 compounds of the CiPA reference set, whose classes are published. Twelve drugs were used to fit two ordinal cut points. Those cuts were then frozen and the remaining sixteen were scored blind.

The comparator is deliberately unflattering to us. Its score is a single number: −log10 of the hERG IC50 divided by the free Cmax. No simulation, no solver, no model. Two ordinal cut points are then fitted to it by the identical protocol, so the two methods are scored on the same footing and neither gets a free threshold.

The reference file carries two published hERG IC50 columns, a dynamic-protocol fit and a static fit, and they disagree by up to seventeen-fold on a single drug. Our harness uses the static column, and that is the column behind every number on this page. Feed the margin the dynamic column instead and it scores 11 of 16 blind and 19 of 28 overall. The baseline beats us by one compound on the numbers we ran, and by two on the ones we did not.

three-class accuracy iPSC-CM simulation hERG-only margin always say “Intermediate”
12 training drugs 8 of 12 8 of 12 4 of 12
16 blind drugs 9 of 16  (0.56) 10 of 16  (0.63) 7 of 16  (0.44)
all 28 17 of 28  (0.61) 18 of 28  (0.64) 11 of 28  (0.39)
Two published numbers and a division beat a CVODE solve per concentration per compound. On this task, on this panel, by this metric.

Read the third column too. A classifier that ignores the compound entirely and always answers “Intermediate” gets 11 of 28. Both real methods clear it, but not by the margin either of us would like, and it is the number most benchmark pages leave out. It comes back in the next section, where it ties us on the one metric we had been treating as a win.

One caveat on “costing nothing”. The margin is free here because someone else already paid for both of its inputs and published them. On a compound of yours it needs a hERG patch-clamp IC50 and a free clinical Cmax, which is an assay and a Phase 1 result. What is free is the arithmetic, not the data.

Where each method wins and loses

Exact accuracy is one metric and it hides the shape of the errors.

iPSC-CM simulation hERG-only margin always “Intermediate” published ORd / qNet
three-class accuracy, 16 blind 0.56 0.63 0.44 not reported
adjacent accuracy, all 28 28 of 28 27 of 28 28 of 28 not reported
High-vs-rest AUC 0.91 (28) 0.89 (28) 1.00 (CiPA validation subset)
Low-vs-rest AUC 0.85 (28) 0.77 (28) 0.89 (CiPA validation subset)
High-risk drugs caught, all 28 6 of 8, 1 false alarm 6 of 8, 2 false alarms 0 of 8 not reported
compounds under-called, all 28 2, none to Low 4, two of them to Low 8, none to Low not reported
head-to-head, all 28 right on 2 the margin misses right on 3 we miss not compared
calls it distributes across 28 20 Intermediate, 7 High, 1 Low 14 Intermediate, 8 High, 6 Low 28 Intermediate
true distribution 11 Intermediate, 9 Low, 8 High

Four things fall out of that table. Two are bad for us, one is good, and one is a claim we made in an earlier draft of this page and have had to withdraw.

Our perfect adjacent accuracy is not a virtue. This is the withdrawn claim. We wrote 28 of 28 in bold and called it a win. Then we added the third column, and a classifier that ignores the compound and always answers “Intermediate” also scores 28 of 28, because a middle answer is never two classes from anything. Our adjacent accuracy is a consequence of hedging, not of modelling. It stays in the table because removing it would be worse, but it should not have been highlighted and it is no longer.

We are badly biased toward the middle. That is the same finding said plainly. The simulation issues “Intermediate” for 20 of 28 compounds and manages a single Low call in the entire panel. The hERG margin distributes its calls far more like the truth. If you want a method that will clear a compound, ours is currently not it.

Where we do win, it is on under-calls. This is the one column that goes our way, and it is the column a safety scientist cares about, because a missed liability costs more than a false alarm. Two compounds are under-called by the simulation, both High drugs dropped to Intermediate, and neither reaches Low. The hERG margin under-calls four, and two of those go all the way to Low: astemizole and risperidone. Astemizole was withdrawn from the market over QT prolongation, its free Cmax is 0.26 nM against a hERG IC50 of 33.3 nM, and a 128-fold margin reads as safe on a spreadsheet. The full read on astemizole is here.

Verapamil is not the win we first described. It is the one drug the hERG margin misses by two classes: clinically Low, called High, because verapamil blocks calcium hard enough to offset its own repolarization delay and a hERG-only score cannot see the offset. But we do not get verapamil right either. We call it Intermediate. The margin fails it by 0.038 log units against a cut point fitted on twelve drugs, and a shift of 0.04 in that cut (about seven percent) erases the error entirely. Head to head, the simulation is right on two compounds the margin misses (astemizole, risperidone) and wrong on three the margin gets (loratadine, nifedipine, nitrendipine). Net, we are one behind. Verapamil is a good illustration of what multichannel simulation is for. It is not evidence that ours works.

Neither of us is the published state of the art. The ORd/qNet column is the CiPA consortium’s own result and its High-vs-rest separation is perfect. Note the label: those two figures are reported on CiPA’s own validation subset, not on all 28, so the column is not on the same footing as ours. The published 1.00 carries a 95% confidence interval of 0.92 to 1. We reproduce the column here rather than omit it.

how to discount these numbers
  1. 28 compounds is a small panel. One drug is 3.6 percentage points. The gap between 17 and 18 of 28 is one compound, and it should not be read as a stable ranking of the two methods.
  2. The AUC columns use all 28. The cut points were fitted on 12 of them. AUC is computed on the continuous score rather than the cuts, so it is not directly contaminated, but the metric was chosen with the data already in view. Treat the blind 16-drug accuracy as the number to trust and the AUCs as supporting.
  3. The ORd/qNet column is quoted from publication, not re-run by us. We did not reproduce it in the same harness, so it is not scored on an identical footing.
  4. Our simulation scores three currents. IKr, ICaL and peak INa. There is no late-sodium current in the model. Ranolazine and mexiletine are both clinically Low, both are late-sodium blockers, and we call both Intermediate. We tested what a late-sodium current would change: adding one recovers mexiletine but not ranolazine, and lowers overall blind accuracy. So the gap explains one of the two misses outright and is only part of the ranolazine story. The model note sets out the gap.
  5. Adjacent accuracy is not an independent metric here. A classifier that always answers “Intermediate” scores 28 of 28 on it. Our 28 of 28 follows from the same hedge, not from the simulation.

A harsher metric on the same data, and a correction

what this section used to say, and why it was wrong

This section previously called the run below “a second independent run” and stated that its numbers were not comparable with the ones above. Both statements were false, and an adversarial review of this page caught them by reading our own source. The harness loads cipa_validation_results.csv, the output of the run in section 02, and uses its risk column directly as a feature. It scores the same 28 drugs, on the same 12/16 split, from the same reference file. And its metric is High-vs-rest AUC, which section 02 already reports. It is one result presented a second way, not a second experiment. We have corrected the framing rather than deleted the section, because the correction is the more useful thing to publish.

What the run does add is a discipline the first one lacks: a locked hold-out, a permutation null, and a pass/fail gate fixed before the scores were seen. Restricting to the 16 held-out drugs, High-vs-rest AUC is 0.9375 for the hERG margin and 0.9062 for our score — the same two numbers section 02 reports on all 28, recomputed on the blind split.

modellocked AUC95% CIgate
hERG-only baseline0.93750.78 – 1.00— it is the baseline
multichannel logistic0.9375fails both gates, p = 0.235
multichannel sum0.8750fails
multichannel + Kernik0.9375fails
Kernik AP score alone0.90620.72 – 1.00fails — does not beat the baseline
Nothing beat hERG-only on CiPA-28. That is the recorded conclusion of that run, and we did not go back and loosen the gate until something passed.

The harness for this run is public at github.com/Perturb-Bio-Inc/cardiac-safety-eval, including the line that reuses the first run’s output. The three-class harness in section 02 is not yet published; if you want it, ask and we will send it.

Two things in that table deserve more than a dash. The multichannel logistic model does not merely tie the baseline. Its permutation p is 0.235, so the harness cannot separate it from shuffled labels at all. And the confidence intervals on the two AUCs overlap almost completely. Sixteen drugs, four of them High, gives 48 pairwise comparisons, so the AUC can only move in steps of about 0.021. The gap between 0.9375 and 0.9062 is one and a half pair swaps. This table cannot rank these methods. It can only say that none of ours cleared a gate set in advance.

One note on the permutation null. The Kernik score reaches p = 0.014 against a null mean of 0.4996, which is well centred. But that candidate fits nothing; it is a fixed column passed through. The fitted candidates on this same benchmark have nulls near 0.37, which the harness README records as a known small-n artefact. Do not read 0.014 as evidence the model is good. Read it as evidence the column is not random.

Three conclusions people draw from this that are wrong

“so simulation is useless”

No. It means simulation is not the cheapest way to produce a risk label. The label was never the expensive part. Verapamil is the standing counterexample: it is precisely the compound where the single-number method fails hardest, and it fails for a reason that only a multichannel model can express.

“so just use the hERG margin”

For triage, often yes, and we will say so on a call. But a margin cannot tell you which current drives the signal, cannot tell you when two mechanisms fit your data equally well, and cannot name the assay that separates them. It answers a different question.

“so their mechanism reads are validated”

Also no, and this is the one we most want you to hold on to. What is benchmarked on this page is the risk label. The mechanism ranking, the thing we sell, has been tested only against phenotypes the same model generated. That shows the limits of what the method resolves. It is not evidence of accuracy against a real recording, and we have not run that test.

what would change our mind

A set of iPSC-CM recordings with the underlying channel pharmacology independently known, scored blind. That is a wet experiment we have not run. Until someone runs it, the mechanism read stays “internally consistent, externally unvalidated”.

Which of these you want

use a hERG margin when

You need to rank a series, you need a number today, and the decision it feeds is “which three of these twelve do we take forward”. It is free, it is fast, it is on this page beating us, and no one needs to be paid for it.

use the CiPA pipeline when

Your next step is a regulatory conversation. It is the published, consortium-backed route, and on the two metrics where a published number exists it is ahead of both of us. No three-class accuracy is published for it, so that comparison is partial. We do not compete with it and we are not qualified for a submission.

use a mechanism read when

You have a repolarization signal you do not understand, a multichannel profile where the currents pull against each other, and a real decision about which experiment to run next. The deliverable is a ranked short list of causes with the distance between them measured, an explicit statement of what your data cannot separate, and one assay that separates it. None of that is a risk class, which is why none of it appears in the tables above.

If a hERG margin already answers your question, use it rather than paying us. A customer who buys the wrong thing complains later, and this company is too small to survive that.

the work these numbers do not measure

Four public CiPA compounds, read end to end: the ranked mechanisms with their scores, what the data cannot separate, the single confirming assay, and the two cases we get wrong. The first read on a public or non-confidential compound is free.

get the numbers

Every per-drug score and class behind this page, one row per compound, plus the two fitted cut points and both confusion matrices:

cipa28-per-drug.csv · all 28 compounds, our score and class, the hERG margin score and class
cipa28-cuts-confusion.json · the fitted cut points and the training and validation confusion matrices for both classifiers

Both are regenerated from the same stored results the tables above quote. If the two disagree, the files are the source of truth and the page is wrong.

Sources and provenance. Compounds, ion-channel block data and torsade classes are the CiPA reference set, taken from the FDA/CiPA repository (AP_simulation/data/newCiPA.csv, branch Model-Validation-2018), and drug categories from Colatsky et al., J Pharmacol Toxicol Methods 2016. The hERG IC50 used throughout is the static fit from Li et al. 2017 (Hill_fitting/data/Li2017_IC50.csv); it is the primary input to both classifiers, not a cross-check, and section 01 gives the result on the dynamic column instead. Simulations use the Kernik-Clancy 2019 human iPSC-CM model (doi:10.1113/JP277724) in Myokit 1.39.2 with CVODE; the model note explains the choice and its cost. The published ORd/qNet AUCs are quoted from the CiPA literature (Li et al. 2018), were not re-run by us, and are reported there on CiPA’s own validation subset rather than on all 28. The three-class comparison is our internal validation run of 2026-07-03. Every accuracy count, the always-Intermediate baseline, the under-call and head-to-head rows, the dynamic-column refit and the verapamil cut-point sensitivity were recomputed from the stored per-drug results on 2026-09-06 and agree with it. The locked-AUC run in section 03 uses a leakage-controlled harness with its own hold-out and permutation null, applied to the same 28 drugs and the same score columns; see the correction at the head of that section. This page was rewritten on 2026-09-06 after an adversarial review found five overclaims in the previous draft. The findings it raised are stated in place rather than quietly fixed.

get in touch
[email protected]