TTwin

A cancer digital twin which is built around the cell population, and not around a risk score.

Cancer prediction models take a patient and compress him into a list of features, then they give back a probability. In this way three things are lost.

A tumour which contains a rare drug-resistant subpopulation, after the averaging it looks the same like a tumour that has none. The tumour burden is disappearing when the features are normalised. And chemotherapy after immunotherapy gives the same feature vector like the opposite order, so the sequence can not be written at all.

PHENOFLUX keeps the population. The state is a distribution over the cell phenotypes with a real mass on it, and every therapy is a function which takes a patient and gives back a patient. So the sequences, the drug holidays, and also the regimens which nobody received, all of them are only function application.

Results

For predicting the relapse it is worse than the standard survival models. On 1,979 METABRIC patients it reached 0.601 concordance, but a random survival forest with the same inputs reached 0.677. On a second cohort the gap stayed also.

But it can do some things which those models can not do. The endocrine therapy operator was written many months before this data was opened. When it is applied on the real patients, it predicted hazard ratio 0.89 for ER-positive disease and 0.98 for ER-negative, and both of them are inside the observed confidence intervals. Also a chemotherapy operator which had the relation with proliferation in the wrong direction was corrected after we fit it on the measured treatment response.

3,697
patients, from four public cohorts
0.601 / 0.677
concordance, the model against a survival forest
−0.21 → +0.11
chemotherapy operator, before and after the calibration