Deep Learning Model Accurately Segments Infarct on Cardiac MR

Automated DL model matches experts on infarct size, with limits in microvascular obstruction


Matthias Schwab, PhD, MSc, BEd
Schwab

Quantifying myocardial infarction on cardiac MR is accurate but time-consuming and subject to inter-reader variability. A new deep learning (DL)-based model shows promise in assisting radiologists with the segmentation and quantification of cardiac MR (CMR) images following myocardial infarction. Research published in Radiology Advances compared the DL-model and human-based measurements, demonstrating the model’s strengths and limitations.

In the study, Matthias Schwab, PhD, MSc, BEd, a postdoctoral researcher at the Medical University of Innsbruck, led a team in developing and evaluating an automated pipeline to assess infarct size and microvascular obstruction (MVO) on late gadolinium enhancement (LGE) CMR images.

“Our deep learning-based algorithm quickly segments myocardial infarction and MVO on clinical images in a fully automated way, without human intervention or pre-processing,” Dr. Schwab said, highlighting the model’s capability.

Automating the assessment of infarct size and MVO on CMR images could improve reproducibility while reducing the time required for infarct quantification after myocardial infarction.

Cascaded Model Trained on LGE CMR Data

To address these challenges, the DL model, a cascaded framework of 2D and 3D convolutional neural networks, was trained on an in-house dataset of 144 manually segmented LGE CMR exams. The training data were part of the Magnetic Resonance Imaging in Acute ST-Elevation Myocardial Infarction (MARINA-STEMI) trial, obtained between 2006 and 2023 at a single center.

Dr. Schwab explained the step-by-step approach. “The first step is segmenting the left ventricle, then zooming into that area, followed by the segmentation of the scars. It's a cascade of three convolutional neural networks,” he said.

For training and evaluation, images included in the training were acquired at baseline, typically within three days of primary percutaneous coronary intervention, and again at four- and 12-month follow-up examinations.

A separate dataset of 152 exams from the same institution was segmented automatically with the DL model and manually by trained doctoral candidate research assistants. The research assistants completed training on a separate dataset of 30 patients and were required to demonstrate acceptable inter-rater reliability before participating in the study.

“In testing, the researchers found good agreement between the manual and automated AI-based calculations,” Dr. Schwab said.

Specifically, the DL method achieved moderate-to-high overlap with manual segmentations, with Dice scores of 64% for infarcts and 82% for MVO. Automated volume measurements closely matched manual assessments, differing by an average of 5 mL for infarcts and 0.6 mL for MVO.

Although MVO was present in only 15% of all patients in the test cohort and achieved an overall Dice score of 82%, the score dropped to 25% among patients with MVO present.

MRI feature

Agreement Strong for Infarct, Weaker for MVO

To qualitatively evaluate the methods, two CMR experts, each blinded to the segmentation method, compared the accuracy of the human and DL-generated segmentations. The expert reviewers had six and 16 years of CMR experience, respectively.

On a per-slice level, the experts were asked to rate the segmentations as optimal, too big, too small, wrong tissue, false negative, positive and true negative. They were also asked to identify which segmentation they preferred overall.

Performance was weaker for MVO detection, with AI-based calculations achieving a sensitivity of 65%. Many MVOs were overlooked or only partially segmented. The experts preferred manual measurements in 55.6% of cases, compared with 11.3% for AI-generated segmentations, while rating the two as equivalent in 33.1% of cases.

The AI model performed considerably better in identifying myocardial scars, missing only two scars across all cases and outperforming human readers in scar identification (2.6% missed vs. 4.2%). Consistent with these findings, experts preferred the AI-based segmentations in 33.4% of cases, compared with 25.1% for human-generated segmentations, while rating the methods as equivalent in 41.5% of cases.

“The experts deemed the DL model better at representing the actual extent of the infarction,” Dr. Schwab said. “We think this preference is due to the consistency of the AI segmentations across slices and patients, which avoided the variability and subjectivity that often characterize manual annotations.”

“Our study demonstrates that the DL model performed on par with or better than human experts for specific conditions, such as scar segmentation,” Dr. Schwab added. “But in other areas, such as MVO, it needs improvement.”

Dr. Schwab said greater volumes of data must be analyzed in future studies. He noted that clinical validation will require multi-center studies with heterogeneous and more diverse populations, as well as broader inclusion of vendors and readers. The training and test datasets were derived from a single institution.

“We will investigate whether these AI segmentations are able to predict specific clinical patient outcomes compared to manual segmentations,” he concluded.

For More Information

Access the Radiology Advances study, “Deep learning pipeline for fully automated myocardial infarct segmentation from clinical cardiac MR scans.”

Read previous RSNA News stories on cardiac imaging: