{"version":1,"type":"story","url":"https://digestai.news/story/deep-learning-perturbation-models-outperform-baselines-under-calibrate","json":"https://digestai.news/story/deep-learning-perturbation-models-outperform-baselines-under-calibrate.json","markdown":"https://digestai.news/story/deep-learning-perturbation-models-outperform-baselines-under-calibrate.md","slug":"deep-learning-perturbation-models-outperform-baselines-under-calibrate","headline":"Deep learning perturbation models outperform baselines under calibrated metrics, study finds","summary":"A study published in Nature Machine Learning introduces a calibration framework showing that deep learning models for genetic perturbation prediction can outperform uninformative baselines when evaluated with well-calibrated metrics. Researchers analyzed 14 datasets and 18 metrics, finding that common benchmarks like mean squared error (MSE) and Pearson correlation of control-referenced deltas are frequently miscalibrated due to control bias and signal dilution. They developed a dynamic range fraction (DRF) metric using positive (interpolated duplicate) and negative (mean baseline) controls to quantify calibration. Weighted metrics (WMSE, weighted R²Δ) and retrieval-based metrics (NIR) showed consistently higher calibration across datasets. Under these calibrated metrics, nine models including scGPT, GEARS, PRESAGE, scLambda, CellFlow, and foundation model embedding probes (Geneformer, ESM2, scGPT, GenePT) mostly outperformed baselines on unseen perturbation and combination tasks. PRESAGE exceeded the additive baseline on 15 of 18 metrics in the Wessels23 dataset, which has greater combinatorial coverage (10.3% vs 0.63% in Norman19). Model performance correlated with biological pathway recovery (ρ=0.76-0.84). The authors caution that conclusions reflect a snapshot of current models and datasets.","keyPoints":["Study introduces dynamic range fraction (DRF) to calibrate 18 metrics across 14 genetic perturbation datasets","Weighted metrics (WMSE, weighted R²Δ) and NIR show consistently higher calibration than MSE and Pearson Δ","Nine deep learning models outperform uninformative baselines under well-calibrated metrics on unseen perturbation and combination tasks"],"whyItMatters":"Recalibrates the debate on genetic perturbation modeling feasibility, showing prior reports of model underperformance stemmed from miscalibrated benchmarks rather than model failure.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":["scGPT","GEARS","PRESAGE","scLambda","CellFlow","Geneformer"],"people":[]},"firstPublishedAt":"2026-10-01T00:00:00Z","updatedAt":"2026-10-01T00:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"Nature Machine Learning","title":"Deep learning perturbation models can outperform baselines on calibrated metrics","url":"https://nature.com/articles/s41587-026-03307-w","publishedAt":"2026-10-01T00:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Deep learning perturbation models outperform baselines under calibrated metrics, study finds\", 1 October 2026, https://digestai.news/story/deep-learning-perturbation-models-outperform-baselines-under-calibrate","publisher":"Digest AI","title":"Deep learning perturbation models outperform baselines under calibrated metrics, study finds","datePublished":"2026-10-01T00:00:00Z","url":"https://digestai.news/story/deep-learning-perturbation-models-outperform-baselines-under-calibrate"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}