PMID 33797598 — A novel deep-learning-based approach for automatic reorientation of 3D...
good_imrad R=1202w / 12¶ | figs=35 Arani
TITLE
[1] 12w A novel deep-learning-based approach for automatic reorientation of 3D cardiac SPECT images
ABSTRACT
[1] 297w Purpose Reconstructed transaxial cardiac SPECT images need to be reoriented into standard short-axis slices for subsequent accurate processing and analysis. We proposed a novel deep-learning-based method for fully automatic reorientation of cardiac SPECT images and evaluated its performance on data from two clinical centers. Methods We used a convolutional neural network to predict the 6 rigid-body transformation parameters and a spatial transformation network was then implemented to apply these parameters on the input images for image reorientation. A novel compound loss function which balanced the parametric similarity and penalized discrepancy of the prediction and training dataset was utilized in the training stage. Data from a set of 322 patients underwent data augmentation to 6440 groups of images for the network training, and a dataset of 52 patients from the same center and 23 patients from another center were used for evaluation. Similarity of the 6 parameters was analyzed between the proposed and the manual methods. Polar maps were generated from the output images and the averaged count values of the 17 segments were computed from polar maps to evaluate the quantitative accuracy of the proposed method. Results All the testing patients achieved automatic reorientation successfully. Linear regression results showed the 6 predicted rigid parameters and the average count value of the 17 segments having good agreement with the reference manual method. No significant difference by paired t-test was noticed between the rigid parameters of our method and the manual method (p > 0.05). Average count values of the 17 segments show a smaller difference of the proposed and manual methods than those between the existing and manual methods. ConclusionThe results strongly indicate the feasibility of our method in accurate automatic cardiac SPECT reorientation. This deep-learning-based reorientation method has great promise for clinical application and warrants further investigation.
INTRO
[1] 209w Single photon emission computed tomography (SPECT) is one of the most commonly used procedures for cardiac imaging, providing information on the relative perfusion of segments of the left ventricular (LV) wall. Along with collection of an EKG signal, it also provides functional information on the segmental heart muscle contraction. Together, these enable diagnosis of many heart diseases [1,2]. The conventional reconstructed SPECT images are transaxial images which are perpendicular to the long axis of the patient, but not perpendicular to the long axis of the left LV. Doing image analysis and diagnosis directly from the transaxial images could lead to regional differences [3] and further cause misjudgment of the heart status [4,5]. To make a diagnostic interpretation from This article is part of the Topical Collection on Advanced Image Analyses (Radiomics and Artificial Intelligence) the cardiac SPECT images clinically, reconstructed images usually need to be reoriented to specific standard planes, i.e., short-axis (SA), horizontal long-axis (HLA), and vertical long-axis (VLA) slices [6,7]. SA slices, which are perpendicular to the LV's long axis, are also necessary for polar map generation to achieve LV quantification and display myocardial perfusion data in a two-dimensional image [8]. Previous studies have demonstrated the incorrect reorientation could result in image artifacts and quantitative error [9,10].
[2] 94w The traditional procedure for cardiac SPECT image reorientation typically requires manual manipulation by the user. The procedure typically consists of the middle transaxial slice of the LV being selected by the user and a line was drawn to delineate the LV's long-axis component (Fig. 1a). This process was repeated in the sagittal plane (Fig. 1b). These two LV's long axes in different planes then defined the coordinate changes from transaxial images to SA images (Fig. 1c). This manual process is time consuming and operator dependent, which leads to the need for an automatic approach.
[3] 191w Automated techniques for reorientating the LV to oblique SA slices have been proposed before and are routinely used in clinical applications [11][12][13]. Two methods start by segmenting or delimitating the LV region using a thresholdbased approach. One approach then used the segmented LV to refine the estimation of the myocardial surface and fit it to an ellipsoid whose long axis is used as the final LV long axis [12]. Another technology tessellated the segmented LV data to triangular plates and the intersections of the normal for each plate on the myocardial surface were collected and fit to a line as the LV long axis [11]. Slomka et al. developed an approach for automatic reorientation by registering the original data to a statistical heart template image [13]. All methods perform well as an alternative to manual reorientation; however, they rely heavily on the overall structural integrity of the LV. Severe myocardial ischemia exhibits large areas of low intensity in cardiac SPECT images and will further affect the automatic long-axis delineation and cause failure of the reorientation [14]. Current commercial software also provides a manual alternative in the event that automated reorientation fails.
[4] 155w Besides SPECT imaging, several automatic cardiac image reorientation methods have been proposed in other imaging modalities. DeKemp et al. [15] proposed an ellipsoid fitting approach for LA estimation in cardiac positron tomography images, and insensitivity to cardiac defects was noticed for this method. Several different attempts have been tried in cardiac MRI. Lelieveldt et al. proposed an automatic SA cardiovascular magnetic resonance image reorientation method using fuzzy implicit surface template [16][17][18]. Jackson et al. proposed a semiautomatic approach for estimation of the reorientation parameters, with user interaction for cardiac chamber segmentations [19]. Lu et al. proposed machine learning-based methods to localize and delineate cardiac anatomies in 3D volume and detect a set of cardiac landmarks for heart reorientation [20]. Due to the image differences between SPECT and MRI, for example, SPECT has poorer spatial resolution and thus less anatomical structure information; these methods may not be feasible for the application in automatic cardiac SPECT reorientation.
[5] 111w Recently, convolutional neural networks (CNN) have been proposed for organ or tissue localization and automatic reorientation in medical images. Hu et al. proposed using deep CNN to automatically localize and segment multiple abdominal organs [21]. Wolterink et al. used a CNN to learn to identify the coronary centerline direction and lumen radius directly from cardiac CT angiography images [22]. Vigneault et al. designed a novel CNN network, i.e., Omega-Net, to achieve automatic twodimensional image orientating relative to a specific position and segmenting the heart structure in cardiac MRI [23]. To the best of our knowledge, there is no study trying to use a CNNrelated method for automatic cardiac SPECT data reorientation.
[6] 85w In this study, we proposed a novel deep-learning-based method for fully automatic reorientation of cardiac SPECT images without the need of prior segmentation for the heart structure. The proposed approach was trained by a group of 6762 datasets (90% for training and 10% for validation) which were generated from 322 subjects, and was evaluated on 52 subjects from the same center and 23 subjects from another center. The reorientation results from manual method as well as an existing method were used as reference for comparison.
RESULTS
[1] 233w The spatial transformer network (STN) [26] was a deeplearning module originally proposed as a general layer which could be included into a standard neural network architecture to provide 2D spatial transformation of the network feature map. It contains three submodules: a localization network to predict the transformation matrix; a grid generator to create a sampling grid, which is a set of points where the input map should be sampled to produce the transformed output; and a sampler to apply the transformation matrix on the input. STN could be used to learn the transformation parameters which were best for the final classification task with the ground truth transformation unknown. It was usually directly implemented on the feature map but not the input original images. In this study, we aimed to learn the transformation from the input transaxial images into the standard SA images, by using CNN modules to predict the 6 rigid-body transformation parameters and using the grid generator and the sampler to apply the rigid-body parameters to generate the reorientated images. The proposed network architecture is shown in Fig. 2. Gradient images were calculated for the original 3dimensional transaxial SPECT images and combined with the transaxial images as the 2 channels' input images (128 × 128 × 128 × 2) of the network. We can get the gradient image through computing the gradient magnitude of each pixel in the image as Eq. 1:
[2] 16w where g x , g y , and g z are the gradients for each direction:
[3] 101w The 6 rigid-body registration parameters were predicted from the CNN layer following the max pooling step and the fully connected layer. The predicted 6 rigid parameters would be separated into translation matrix M ([t x , t y , t z ]) and rotation matrix R ([β, α, γ]) and the rotation matrix R was then converted from Euler angle to World Coordinate System angle to fit the input format of the following 3-dimensional grid generator and the sampler. Due to the convention of elastix, the Euler angles are ordered as YXZ; thus, the rotation matrix could be converted as follows:
[4] 32w The processed rotation matrix R and translation matrix M were re-combined to [R M T ] and served as the input to transform the input transaxial images to the reoriented output images.
[5] 126w Figure 3 shows six sample patients with the 3 orthogonal views of the input reconstructed images, reference reoriented images, and the predicted output reoriented images. The heart structures of patients bv966 and bv1083 have reduced localization areas but the LV contours are generally complete while patients bv844 and bv929 suffer severely from activity decrease which leads to obviously incomplete LV structure. Also note that patient bv966 has been interpreted as having inferolateral ischemia but it has not caused visually obvious structural incompleteness. Patients bv1162 and bv1294 have normal hearts with visually clear heart structure. Visual assessment shows that the proposed methods can achieve automatic reorientation successfully by demonstrating similar performance with the manual method, for both normal hearts and hearts with mild or severe low-uptake defects.
[6] 126w Quantitative results for the 6 rigid parameters are shown in Fig. 4 and Table 3. Linear regression analysis result in Fig. 4 shows that the predicted 6 rigid parameters having excellent agreement to the manual method with the coefficient of determination (R 2 ) are 0.928, 0.958, and 0.994 for the three translation parameters, while 0.960, 0.973, and 0.970 for the three rotation parameters respectively. Table 3 shows the mean and 4 Linear regression analysis result of the predicted 6 rigid-body parameters from the proposed method compared with the manual method the STD values from the proposed automatic and manual methods and the corresponding paired t-test result. No statistical difference was found for the parameters between the proposed method and the manual method (p > 0.05).
[7] 154w Figure 5 shows 8 sample polar maps generated from the reoriented SA slices using the reference manual method and the proposed automatic method. The 17-segment plot of the relative difference value calculated from Eq. 8 is also shown in Fig. 5 accordingly. Among the 8 exampled patients, four have obvious myocardium regions of low uptake (bv760, 844, 929, and 966), two have high uptake due to overlapping of the LV myocardium and the abdominal organs in the inferior region (bv813 and 1083), and the rest have normal LV with uniform uptake (bv1162 and 1294). Three of the patients decreased areas of uptake were interpreted as having severe myocardial activity decrease (bv760, 844, and 929). Polar maps are visually similar between the manual and proposed automatic methods. The mean RD values for the 8 patients BV760, BV844, BV929, BV966, BV813, BV1083, BV1162, and BV1294 are 4.57%, 3.72%, 2.66%, 2.10%, 5.31%, 4.46%, 1.72%, and 1.78% respectively.
[8] 95w Linear regression analysis result of the average count value pairs from the total 884 segments is shown in Fig. 6 with good agreement (R 2 = 0.9926) between the manual method and the Table 4 The RD-seg value of the 17 segments among 52 patients from center A S e g m e n t# 1 2 3 4 5 6 7 8 9 RD-seg 4.93% 3.81% 2.96% 2.90% 3.77% 5.75% 1.86% 1.14% 1.20% Segment # 10 11 12 13 14 15 16 17 Mean(RD-all) RD-seg 0.97% 1.18% 1.20% 1.90% 2.84% 1.39% 1.61% 2.66% 2.47%
[9] 25w proposed method. RD-seg and RD-all results are shown in Table 4. The RD-seg value ranges from 0.97 to 5.75% with the overall average of 2.47%.
[10] 81w Both the existing and the proposed methods successfully achieved automatic heart reorientation for the 23 patients. But we noticed that for the heart with a generally normal heart structure, both methods performed well (FZ01 in Fig. 7) while for heart with incomplete or unclear structure, the existing method may sometimes lead to imperfect result (FZ12 in Fig. 7). Among the 23 patients, we also noticed that both the two methods induced incorrect reorientation result in one patient (FZ16 in Fig. 7).
[11] 132w Figure 8 shows five sample polar maps and their corresponding 17-segment plots of the RD value. Two out of the five patients have normal hearts with generally uniform intensity distribution (F01 and F14). F12 and F23 are two patients with obvious heart ischemia, which is reflected in the polar maps by the apparent low-uptake region. F16 is the one with a normal heart but both our proposed method and the existing method did not perform perfectly. From the polar maps, no obvious difference is noticed visually between the different methods except in F12 with the existing method showing slight variation. The mean RD values for the 5 patients are 0.66%, 3.23%, 4.92%, 8.47%, and 3.91% for the proposed method, and are 3.32%, 1.48%, 22.8%, 10.88%, and 4.51%, for the existing method respectively.
[12] 81w Figure 9 shows the linear regression analysis result of the total 391 segments using the manual method as the standard reference. Both methods show good agreement with the manual method while the proposed method demonstrates better fitting result as well as the higher R 2 value than the existing method (0.9823 > 0.9627). RD-seg and RD-all results of the two methods are shown in Table 5. The RD-all values of the proposed and the existing methods are 3.94% and 6.32% respectively.
DISCUSS
[1] 178w Incorrect reorientation of the LV in cardiac SPECT images can cause visual misdiagnosis and quantification errors in the subsequent polar map analysis and the quantitative evaluation. Manual reorientation should be conducted to align the LV to the standard SA direction and could be used as a gold standard. However, manual processing suffers from relatively FZ01 Transaxial Coronal Sagital SA HLA VLA SA HLA VLA SA HLA VLA Transaxial Coronal Sagital SA HLA VLA SA HLA VLA SA HLA VLA FZ12 FZ16 Transaxial Coronal Sagital SA HLA VLA SA HLA VLA SA HLA VLA d e t c u r t s n o c e R Images e c n e r e f e R s e g a m I A S Predicted s e g a m I t u p t u O Existing t u p t u O d o h t e M Fig. 7 Three sample patients with 3 orthogonal orientations of the input reconstructed slices, reference reoriented slices, the predicted output reoriented slices, and the existing method output slices
[2] 74w high intra-and inter-observer variability and results in unstable performance [12]. Although automatic reorientation methods have been proposed before and applied in commercial software, they rely heavily on the structural integrity of the heart and are prone to failure in the event of a major myocardial ischemia. We proposed a deep-learning-based automatic reorientation method to study the rigid-body motion mapping between the original transaxial images and the SA images, and predict the rigid-body transformation parameters.
[3] 403w All the 52 patients in our test group of center A and the 23 patients of center B achieved reorientation successfully (100%) even with mild to severe myocardial ischemia (bv760, 844, and 929 on Fig. 5 and FZ12 on Fig. 7) or significant overlapping of the heart and abdominal organs (bv813 and 1083 on Fig. 5). This compares very favorably to the previous studies where 394 of 400 cases (98.5%) [12] and 116 of 124 cases (93.5%) [11] were reported to achieved successful automatic reorientation. Failure in previous studies was mostly described as (1) failure to isolate the LV; (2) the presence of substantial abdominal organ components in the segmented LV image; or (3) the reorientation angles differing by more than 45°from the manually obtained ones, while these failures did not happen in our method. Although slight reorientation angle misalignment might happen when using the testing patient of center B (FZ16), the angle difference was around 20°which is much smaller than that of 45°. When we used the data from the same center for training and testing, results in Fig. 4 and Table 3 showed good correlation between the predicted parameters of our method with the manual method and no significant difference (p > 0.05) was noticed. Visual assessment from the SPECT images (Fig. 3) and the polar maps (Fig. 5) also indicates the good performance of our method. RD analysis results in Figs. 5 and 6 and Table 4 show good correlation and small RD-all of 2.47% value between the manual and our automatic methods. If we changed to use different centers' data for training and testing, i.e., data from center A to train the network and data from center B for testing, we noticed our proposed method 5). Average count values of the 17 segments from the proposed method show better fitting with the manual method (Fig. 8) and the RD-all result of the proposed method is also smaller than that of the existing method, which indicates the potential better effectiveness of the proposed method. Also notice that the testing data from center A showed smaller RD-all than that from center B. Although the good results from center B demonstrate the effectiveness of our trained network for handling the multicenter data, the higher RD-all value and occasionally happening imperfect reorientation results implied that the generalization ability of our method to process data from different centers' data needs to be further improved.
[4] 123w Compared to the previous automatic methods, our proposed method does not need extra pre-processing for the input SPECT images to extract the LV structure as the deep network would learn from the whole image. While for the previous methods, image thresholding and LV structure extraction may be required as additional implementations before the reorientation process starts. The characteristics of deep learning network also lead to its perfect reproducibility for each execution as compared to the manual operation, i.e., for a specific input image, the output would be exactly same for different operations. And besides cardiac SPECT imaging, this method could be extended to be implemented in cardiac PET, MRI, or CT images and we would further evaluate its efficiency in the future study.
[5] 105w A limitation of our study is that we used only 322 patients with data augmentation for the network training, and data augmentation is not a true substitute for additional clinical data. More clinical patients included in the future study would further improve the generalizability of our network and the accuracy of the image reorientation process. Furthermore, image data collected from different clinical center are essential to investigate and improve the multi-center performance [28]. To solve this problem, a multi-center effect correction method directly applied in the network feature map is currently under investigation in our group and its efficacy would be evaluated for our network.
CONCL
[1] 76w We proposed a deep-learning-based approach for automatic LV reorientation of cardiac SPECT images. The deep network learns from the reconstructed transverse slices without any pre-process to predict a group of rigid-body transformation parameters for the reorientation. Preliminary results showed the proposed method operates successfully on all test patients and agrees well with the results from manual reorientation. Thus, our deep-learning approach could be potentially served as an automatic way of LV reorientation for cardiac SPECT images.
METHODS
[1] 148w The proposed cardiac SPECT image automatic reorientation method consisted of a training step for the estimation of the parameters and a reorientation step which made use of these parameters. For a given pair of the reconstructed transaxial images and the manual reoriented SA images, the transformation between them could be treated as a rigid-body registration and these rigid-body parameters could easily be manually calculated and used as the learning-based target. A dataset of 6762 original transaxial images, SA images, and the rigid parameters which were collected from the same clinical center (center A) were used to train and validate the designed deep learning network which were used for the subsequent cardiac SPECT automatic reorientation. A dataset of 52 patients' transaxial images collected from center A and 23 patients' transaxial images collected from another center (center B) were processed with the aforementioned trained network for further analysis and verification.
[2] 71w We retrospectively analyzed 374 anonymized patients who underwent Tc-99m-sestamibi stress SPECT/CT scan at University of Massachusetts Medical School (center A) and 23 anonymized patients who underwent Tc-99m-sestamibi rest SPECT/CT scan at The First Affiliated Hospital of Fujian Medical University (center B). Written consent was collected from all the patients and the study was approved by the local institutional review board. Acquisition protocols of the two centers are listed in Table 1.
[3] 207w The reconstructed transaxial images of all patients were manually reoriented by a medical physicist with extensive experience to SA images using IDL software provided by University of Massachusetts Medical School [24]. Rigid-body registration was conducted between the transaxial and SA images to calculate the 6 rigid registration parameters, using the elastix registration program [25]. The 6 parameters contained 3 degrees of translation ([t x , t y , t z ]) in the unit of pixel and 3 degrees of rotation ([β, α, γ]) in radian, with the image center as the center-of-rotation. The 374 patients were separated into two groups with group 1 of 322 patients used for net training and validation, while group 2 of 52 patients for data testing. The minimum, maximum, mean, and standard deviation (STD) of the 6 parameters of group 1 are showed in Table 2. We noticed that parameters for rotations have a range of 1.511, 1.867, and 1.850 for β, α, and γ, respectively, which are all around or more than π/2, while the parameters of the translations have a range of ~30 pixels. This data would vary greatly from patient to patient and the purpose of our work was to find the appropriately specific parameters for each patient.
[4] 61w To enlarge the current dataset, we randomly generated 6440 sets of the 6 parameters from truncated Gaussian distributions based on the mean and the standard deviation of parameters as shown in Table 2. The total 6440 sets of data were divided into 322 groups with 20 sets per group. For each SA image of the 322 patients, an inverse transformation was
[5] 88w The reconstructed transaxial slices of all the patients achieved automatic reorientation by the Siemens Syngo MI applications for SPECT cardiology imaging which is an already existing automatic application. Manual reorientation of these 23 transaxial slices to SA slices was also done by a medical physician with extensive experience using the IDL software. Thus, for center B, a group of 23 patients with input transaxial images, manual reorientated images, and auto-reorientated images using existing software were prepared and would be used for the further validation and multi-center effectiveness evaluation.
[6] 78w We proposed a compound loss function incorporating both the effectiveness of the loss between the 6 rigid parameters and the Dice loss functions of the predicted output images with the referenced SA images to supervise our network. L1 loss was used as the metric for the rigid parameters but computed for the translation (L-trans) and rotation (L-rot) parameters separately due to their differing by an order of magnitude. The L1 loss for the parameters was defined as follows:
[7] 44w where P i and R i are the parameter of the prediction and reference for the i th parameter, and n indicates the number of parameters, i.e., n = 3 for the translation and rotation matrices respectively. The parameter loss (L-par) was calculated as:
[8] 24w where μ is an empirical parameter which we took as 10 for our study to balance the loss between the rotation and translation parameters.
[9] 139w To further optimize the network, we combined image loss with the parameter loss. Parameter loss could give the network a better initial start direction about the training and the image loss could directly reflect the final difference between the prediction and the truth. These two losses could constrain and mutually help the training converge faster and better. The predicted output images and the reference images were binarized using the value 1 as the threshold to extract as much of the image contour as possible, i.e., for these two images, pixels with image intensity larger than or equal to 1 were set to the value of 1, while pixels with intensity smaller than 1 were set as 0 to make the original images into binary images. Dice loss was computed between the binary images and used as image loss (L-img).
[10] 133w where B P and B R are the predicted and reference binary images respectively and V() indicates the volume of the image. The previously mentioned 6762 groups of datasets were used to train the proposed network with 6086 groups (90%) used as the training set, while the remaining 676 groups (10%) were used as the validation set to monitor the training progress. The proposed network was implemented using PyTorch on a NVIDIA Titan V GPU with 12 GB memory and the Adam gradient descent optimizer was employed for training. The whole 3-dimensional image was used as a patch for input and batch size was set as 6. Total epochs for training were 65 with first 30 epochs using L-par as the metric and the remaining 35 epochs using a compound loss (L-com) metric:
[11] 30w where λ is an empirical parameter with a value of 3 used for our study to adjust the value of two losses into similar levels for balancing the training losses.
[12] 185w Once the network has been finished training, it could automatically process the input images and generate the prediction of the rigid registration parameters as well as the predicted reoriented images. The 52 patients in group 2 of center A and the 23 patients from center B were used to evaluate the trained network. For each patient, predicted rigid registration parameters were generated and the slices were reoriented by the STN module using these parameters. The input transaxial images were also reoriented by the STN module using the estimated 6 rigid registration parameters from elastix to generate a new group of SA images, which was used to replace the original referenced manual SA images and to ensure the same processing between the predicted and the referenced images. The predicted rigid-body parameter sets were compared with the value calculated from the manual reoriented references using linear regression analysis on the center A data. The average values of the 52 patients for each rigid parameter were computed, and paired t-test was conducted to evaluate if the differences between the proposed automatic method and manual method are statistically significant.
[13] 126w We further evaluate the performance of the method directly on the reoriented cardiac images and quantitative results. Polar maps were generated from the referenced SA images, the predicted output images, and the output images of the existing method from center B using a IDL software provided by University of Massachusetts Medical School. The 17-segment analysis was performed on the generated polar maps based on the American Heart Association guidelines [27]. Average count value (I) was computed for each segment and the relative difference (RD) was calculated for Fig. 2 The proposed deep learning network architecture. Arrow from the input to the STN indicates applying transformation on input images by the STN module the proposed and automated methods compared with using the manual method data as references:
[14] 23w where I P/A is the I value from the proposed method or the automated method and I R is from the manual method.
[15] 116w Linear regression analysis was conducted on the I values for all the segments of the testing patients, i.e., 17 × 52 resulting in a total of 884 pairs of value from center A and 17 × 23 resulting in a total of 391 pairs of value from center B. For each segment, the RD value was averaged (RD-seg) from all the patients in each center and each method. We also averaged the RD-seg value from all the 17 segments (RD-all) and used this as an overall evaluation metric. show diminished uptake in portions of the LV wall to varying extent while patients bv1162 and bv1294 have generally normal LV structure Eur J Nucl Med Mol Imaging