[1]
42w
This article has been accepted for publication and undergone full peer review but has not been through the copyediting, typesetting, pagination and proofreading process which may lead to differences between this version and the Version of Record. Please cite this article as
[1]
220w
Deep learning-based computer vision methods have recently made remarkable breakthroughs in the analysis and classification of cancer pathology images. However, there has been relatively little investigation of the utility of deep neural networks to synthesize medical images. In this study, we evaluated the efficacy of generative adversarial networks (GANs) to synthesize high resolution pathology images of ten histological types of cancer, including five cancer types from The Cancer Genome Atlas (TCGA) and the five major histological subtypes of ovarian carcinoma. The quality of these images was assessed using a comprehensive survey of board-certified pathologists (n = 9) and pathology trainees (n = 6). Our results show that the real and synthetic images are classified by histotype with comparable accuracies, and the synthetic images are visually indistinguishable from real images. Furthermore, we trained deep convolutional neural networks (CNNs) to diagnose the different cancer types and determined that the synthetic images perform as well as additional real images when used to supplement a small training set. These findings have important applications in proficiency testing of medical practitioners and quality assurance in clinical laboratories. Furthermore, training of computeraided diagnostic systems can benefit from synthetic images where labeled datasets are limited (e.g., rare cancers). We have created a publicly available website where clinicians and researchers can attempt questions from the image survey at http://gan.aimlab.ca/.
[1]
114w
Applications of deep learning to pathology images have shown great potential in a range of tasks, including identifying cancer [1][2][3], predicting patient outcomes [4][5][6], and classifying genomic driver mutations [7] and molecular subgroups [8] solely from images. These successes have exclusively used discriminative machine learning methods, which map data (such as images, text, or speech) to a class label, while generative methods, which model the underlying probability distribution of a data set and can synthesize examples of that data, have remained relatively unexplored. We were motivated to study generative modelling in pathology due to its numerous potential uses in education, clinical quality assurance, improving deep learning classifiers, and digital image processing (e.g. nuclei segmentation).
[2]
140w
Generative adversarial networks (GANs) are a recently developed technique that have had great success in synthesizing high resolution, realistic images [9]. GANs consist of a generator network (analogous to a counterfeiter), which creates synthetic images, and a discriminator network (the detector), which takes as input the synthetic images, as well as a set of real training images, and attempts to determine which are real and synthetic. The generator only receives as feedback whether the discriminator is fooled by the synthetic images and uses this to adjust its parameters; importantly, the generator does not at any point in training have direct access to the real images. Since the GAN concept was first published in 2014, there have been numerous advances in training methods, including the integration of deep convolutional architectures [10], and the development of the progressive GAN training method [11].
[3]
30w
Prior medical applications of GANs have included the synthesis of pathologic images of breast cancer [12], gliomas [13], and cervical dysplasia [14], as well as images of macular degeneration [15],
[4]
9w
This article is protected by copyright. All rights reserved.
[5]
29w
dermatologic conditions [16], and several types of radiographic modalities [17], [18], [19]. However, the images generated in these previous works were generally quite limited in size and low resolution.
[6]
14w
There are several factors motivating our interest in generative modelling for medical images [20].
[7]
123w
Synthetically generated images would have significant educational value [21], in that they could provide a near endless source of novel examples of rare pathologies for teaching and proficiency testing, as well as avoiding confidentiality issues relating to including patient images in publicly available documents [22]. Furthermore, these images could address a major issue in training deep neural networks for medical applications-the challenge in obtaining a sufficient amount of annotated training data for the model to capture high level features and prevent overfitting [23]. Data compilation and annotation for medical applications typically requires the involvement of highly trained experts, and enriching data sets with synthetic images can leverage the work of these experts to maximize training material and decrease overall time and cost requirements.
[8]
155w
We hypothesized that using GANs we could generate realistic-looking images that are adequately classifiable by pathologists, and can be utilized for educational purposes as well as leveraged to improve histopathology classifier performance, where limited datasets are available. Our results show that high quality tissue images as large as 1024 x 1024 pixels can be generated by GANs across five different cancer types as well as the five histotypes of ovarian carcinoma, and that expert pathologists cannot differentiate them from real images. Furthermore, we show that, by leveraging this data, we can achieve performance improvements in deep learning models to classify cancer subtypes where limited data This article is protected by copyright. All rights reserved. might be available. In aggregate, our work presents new opportunities in leveraging GANs to simulate realistic pathology images that can be utilized for educational and quality assurance purposes beyond their commonly advocated use for data augmentation in the machine learning community.
[1]
105w
Whole slide images (WSIs) were acquired from The Cancer Genome Atlas Genomic Data Commons (GDC) Portal (https://portal.gdc.cancer.gov/) [32]. FFPE slides for ovarian cancer cases were acquired from the OVCARE archive and were scanned at 40x magnification on a Phillips IntelliSite Ultra Fast Scanner. Representative regions of tumor were annotated by either a board-certified pathologist, or a pathology resident under supervision by a board-certified pathologist. Following annotation, the WSIs were tessellated with stride=1024 (no overlap between patches) and smaller image patches of 1024x1024 pixels were extracted in TIFF format for training (supplementary material, Table S1). During training, random horizontal mirroring was applied for further image augmentation.
[2]
152w
Several types of cancer were used from the TCGA archive in order to represent a wide range of distinctive morphologic characteristics: low-grade glioma (333 slides; 71,043 image patches), liver hepatocellular carcinoma (78 slides; 59,894 images), lung squamous cell carcinoma (72 slides; 29,422 This article is protected by copyright. All rights reserved. images), renal clear cell carcinoma (104 slides, 78,277 images), and papillary thyroid carcinoma (81 slides, 16,069 images). These five cancer types were chosen as they had a large number of slides available in the TCGA archives and were expected to be classifiable with reasonably high accuracy by general pathologists based solely on a small image patch. The slides used from the OVCARE archive included the five histotypes of ovarian carcinoma as follows: high-grade serous (165 slides; 75,861 images), low-grade serous (39 slides; 27,903 images), endometrioid (57 slides; 22,557 images), clear cell (56 slides; 28,246 images), and mucinous (37 slides; 19,000 images).
[3]
192w
In order to represent a wide range of distinctive morphologic characteristics, 668 whole slide images (WSIs) corresponding to five cancer histotypes were taken from The Cancer Genome Atlas (TCGA) archive [32] with the following breakdown: low-grade glioma (LGG; 333 slides), hepatocellular carcinoma (HCC; 78 slides), lung squamous cell carcinoma (LSCC; 72 slides), renal clear cell carcinoma (RCC; 104 slides), and papillary thyroid carcinoma (PTC; 81 slides). These five cancer types were selected based on their being relatively visually distinct and diagnosable based on a very small image. Furthermore, to show the utility of the methods in synthesizing pathology images for subtypes with subtle morphologic differences, we constructed a dataset of 354 WSIs corresponding to the five main histotypes of ovarian carcinoma from the BC Ovarian Cancer Research Program (OVCARE) with the following breakdown: high-grade serous (HGSC; 165 slides), low-grade serous (LGSC; 39 slides), endometrioid (ENC; 57 slides), clear cell (CCC; 56 slides), and mucinous (MUC; 37 slides). Following annotation of representative regions of tumor, smaller image patches of 1024x1024 pixels (at 40x optical zoom) were extracted from the WSIs and these were used to train various types of generative models (Figure 1).
[4]
103w
Due to their lack of an objective performance function, it is challenging to directly evaluate the performance of GANs [36]. A common test is to have human raters evaluate whether the images are real or synthetic. Using a survey consisting of a mixed set of real and synthetic images, we asked pathologists to classify each image as one of the five histotypes, rate whether they consider it to be of sufficient quality to enable histologic classification, and to determine if the image is real or synthetic. This is similar to the approach that was used previously for the assessment of retinal images [15].
[5]
9w
This article is protected by copyright. All rights reserved.
[6]
315w
We created two surveys, the first of which was for the GANs trained on the TCGA slides of five different cancers. Four board-certified pathologists completed this survey (experience = 1-10 years). When asked to assign histotype (Q1), there was no statistically significant difference in the accuracy of the pathologists (compared to the true labels) when diagnoses were made from synthetic versus real images. In fact, pathologists performed slightly better on the synthetic images, with a median Cohen's kappa of 0.77 (compared to the true diagnosis) for classifying the real images, compared to 0.78 for the synthetic images, supporting the equivalency of real and synthetic images (Fisher's exact test p = 1.0; Table 1 and supplementary material, Table S5). The most challenging histotypes to diagnose were HCC and LSCC, which had substantially lower mean classification accuracies than the other three types (supplementary material, Table S6 and Figure S5). There was no major difference in the pathologist assessment of image quality (Q2) between real and synthetic images (median 88% compared to 90% considered adequate for diagnosis, respectively; Fisher's exact test p = 0.82). Finally, the pathologists did not demonstrate an ability to reliably distinguish the real from synthetic images (Q3), with a median performance of 54% (Fisher's exact test p = 1.0), which is only slightly better than guessing with 50% chance. Analysis of inter-rater agreement demonstrated strong concordance (though lower than the concordance between a rater's diagnosis and the true diagnosis) for histotype classification, with a Fleiss' kappa [37] of 0.707, while for the other two questions there was no agreement (Fleiss' kappa=-0.030 for image quality and -0.025 for real vs. synthetic; supplementary material, Table S7). These results suggest that the individual pathologists did better in diagnosing the histotypes (median Cohen's kappa of 0.77 and 0.78 for real and synthetic images), while their agreement amongst themselves in diagnosing individual samples (Fleiss' kappa of 0.707) was lower.
[7]
9w
This article is protected by copyright. All rights reserved.
[8]
213w
The second survey was for the GANs trained on the OVCARE slides, which included the five histotypes of ovarian carcinoma. Five board-certified pathologists in different geographic regions completed this survey (Table 2 and supplementary material, Table S8), all of whom had subspecialty expertise in gynecological pathology (experience = 5-30 years). For the first question of diagnosing the histotype, as expected the pathologists had lower overall accuracy than for the TCGA images but again performed slightly better on the synthetic images, with a median kappa of 0.66 compared to 0.69 for identifying the histotypes for the real and synthetic images, respectively (Fisher's exact test p = 1.0). The most challenging histotypes to diagnose were ENC, HGSC, and LGSC, while a higher proportion of MUC and CCC were correctly diagnosed as such (supplementary material, Table S6 and Figure S5). Overall evaluation of the image quality for histologic assessment revealed that the synthetic images had higher quality compared to real images (median 91% versus 81%; respectively), suggesting that GANs generated images that contain more histological distinctive features that aid the pathologist in better subtype identification. As in the first survey, the pathologists did not demonstrate an ability to reliably distinguish real from synthetic images, with a median accuracy of 54% (Fisher's exact test p = 0.88).
[9]
36w
Similarly, there was high inter-rater agreement for histotype classification (Fleiss' kappa = 0.601) and minimal agreement for ratings of image quality (Fleiss kappa=0.234) or distinction of real from synthetic images (Fleiss' kappa=0.153; supplementary material, Table S7).
[10]
104w
This article is protected by copyright. All rights reserved. participant). Overall, these results were similar to those of the gynecological specialty pathologists, but with the expected lower classification accuracy for the histological subtypes. The trainees had a median classification kappa of 0.58 on the real images and 0.62 on the synthetic images, considered almost identical proportions of images to be high quality (median 90% of real images versus 92% of synthetic images), and had a median accuracy of 0.56 for distinguishing real from fake images. As expected, there was a trend towards better accuracy for classifying the histological subtypes for the more experienced trainees.
[11]
102w
To further evaluate the quality of the synthetic images, as well as establish their value in building deep learning models for histotype classification where limited data is available, we evaluated whether the addition of synthetic images improved deep learning classifier accuracy similarly to the improvement that one would see with the addition of real images. We used a set of 52 additional ovarian cancer WSIs S10 and S11). More specifically, we achieved median AUC of 0.897 for the baseline augmented with real data and median AUC of 0.917 for the baseline augmented with synthetic images (Figure 3 and supplementary material, Figure S6).
[12]
126w
We repeated the same experiments with a set of additional 108 WSIs corresponding to various cancer histotypes (22 LGG, 20 HCC, 20 LSCC, 20 RCC, and 26 PTC) that were not included in the TCGA GAN training. While the median AUC of the classifier for the baseline set was 0.9878, the performance slightly improved by adding more real (median AUC = 0.9879) or synthetic (median AUC = 0.9924) images to the baseline training set (Figure 3; supplementary material, Tables S10 and S11). While the improvement in performance for this data set was smaller than that for the ovarian carcinoma classifier, this is an easier classification task and the baseline performance was already very high, making it less likely that a large effect size could be observed.
[13]
9w
This article is protected by copyright. All rights reserved.
[14]
95w
It should be noted that the performance of the OVCARE and TCGA classifiers improved with the addition of either real or synthetic data to the baseline set. Similar to the assessment of image qualities by pathologists, although statistically not significant, augmentation with synthetic images led to more performance improvements compared to the addition of real images. In aggregate, our results show that synthetic images generated by GANs perform at least similar to (if not better than) real images when used for training deep learning classifiers, suggesting that synthetic images have similar quality as real images.
[15]
9w
This article is protected by copyright. All rights reserved.
[1]
213w
We present a flexible framework that can synthesize high quality histopathology images for a wide range of tissue types. To our knowledge these are the highest resolution synthetic pathology images that have been generated in the literature, and represent a significant advance in the use of generative deep learning in medicine. A set of pathologists from a range of subspecialties were able to classify the synthetic images based on a single very small patch, rated the image quality highly, and had no ability to reliably distinguish between real and synthetic images. We further demonstrate that the synthetic images are objectively useful as a form of data augmentation for classifier training, as enriching our training dataset with synthetic images improves the accuracy of neural networks trained to classify cancer histotypes. In comparison with other generative machine learning models, the GAN framework performed significantly better and generated qualitatively sharper and more realistic images. Future work is needed to increase the size and resolution of the generated images, as well as to investigate other technical variations within the GAN framework. Furthermore, with the increasing use of digital pathology in clinical practice [38], generative models can be used for various digital image processing applications [39,40,41] and applied to forms of semi-supervised learning such as anomaly detection [42].
[2]
202w
Our use of two data sets demonstrates the breadth of input images that can be incorporated into our pipeline. The survey of images from GANs trained on TCGA slides was designed to be a straightforward classification task of five distinct cancer types, while the differentiation of the five histotypes of ovarian carcinoma is a challenging morphologic classification task for general pathologists, yet has evolved to become highly reproducible amongst experts [43,44]. The accurate diagnosis of ovarian carcinoma is of This article is protected by copyright. All rights reserved. critical importance in guiding therapy, as each ovarian carcinoma type has a distinct biomarker profile and natural history [45,46]. The fact that the synthetic images could be reliably classified supports our assertion that they are of diagnostic quality, which in our view implies that they are usable for directly clinically relevant tasks. While the classification accuracy for ovarian carcinoma in our survey was lower than that reported in previous studies [44], the task that we presented to the reference pathologiststhe classification of cancer histotype based on a single small high-magnification image patch-was significantly more difficult than what is required in pathology practice, in which the whole slide or multiple slides are available for analysis.
[3]
205w
Given the desire to use GANs to supplement datasets for rare tumors, an important factor to consider is the number of images that are required for GAN training. We have trained GANs using varying numbers of images, in part due to inherent limitations in the size of our training set and the need for manual annotation of the images. We found that, overall, the minimum number of images needed to effectively train a GAN is approximately 8000, and that lower numbers than this result in an increased amount of artifact in the synthetic images. Despite the well described challenges in training GANs [36,47], we overall found that training proceeded very smoothly with the progressive GAN method, and very rarely experienced issues with training stability. In particular, we did not observe significant mode collapse, a common failure point in GAN training, in which the generator converges to output many copies of the same image. Some artifacts were encountered (see supplementary material Figure S7 for examples), including grid-like linear streaks in the images, lack of separation of nuclei in closely juxtaposed cells, in some cases forming linear chains of multiple nuclei, and, rarely and only with multi-label conditional This article is protected by copyright. All rights reserved.
[4]
158w
GANs, images that were evenly split between two clearly different histotypes. Importantly, images rated of insufficient quality for interpretation were equally frequent amongst the synthetic and real images, demonstrating that significant artifacts were not generated at sufficient frequency to interfere with interpretation of the images. Situated within the broader context of the implementation of artificial intelligence in the health care system, generative methods can address key concerns regarding data sharing and quality assurance [48]. The current era of "big data" in medicine has the potential to significantly improve patient safety, but concurrently raises major privacy concerns, particularly in the context of data sharing between institutions and the possibility that combining multiple anonymized datasets can allow for reidentification [22]. As the generator network in a GAN has no direct access to the patient images, these models, in combination with differential privacy strategies [49], can capture the features of a dataset yet be transferred with minimal risk to patient privacy.
[5]
292w
While the initial studies applying deep learning to pathology used data sets of several hundred slides [1,2], more recent work has suggested that upwards of 10,000 slides may be needed to develop and validate a clinical support system that is capable of recognizing the full breadth of variability within pathology images [50]. The resources necessary to assemble and securely store a dataset of this magnitude are a significant barrier for many institutions that would otherwise be interested in validating and implementing such systems, and GANs provide the ability to leverage available data and expand access to the technology. They would make it possible to run regular system quality assurance using synthetic images, at a fraction of the cost of manually curating cases to test the system, and ensure that the performance characteristics are acceptable day to day. This same approach could be applied to proficiency testing of individual practitioners in morphologybased specialties. It is well recognized that practitioners can fail to reach accepted standards of practice for a variety of reasons, including related to cognitive skills [51]. While a number of strategies exist for physician performance assessment, they are inconsistently applied and may suffer from high cost and subjectivity (such as oral examinations or performance evaluations). Examinations to test diagnostic proficiency or cognitive ability, for example, are offered but the labor and cost associated with their production means they are only infrequently offered and then only to large cohorts. Proficiency testing on synthetic images could be done independent of a large cohort, with cases regularly refreshed, and at a fraction of the cost, with the output being objective and easily compared to the results of other practitioners. Such a system could also be used in education, to monitor progress through residency training.
[6]
79w
In summary, generative adversarial networks can synthesize cancer pathology images that are indistinguishable from real images to expert observers and are useful as training data for other machine learning algorithms. The synthetic images have numerous applications, ranging from those that are directly and immediately clinically relevant, to others that are more research-focused and technical in nature. We are therefore optimistic that advances in generative modelling in medicine are an important This article is protected by copyright. All rights reserved.
[7]
147w
All images 0.54 (0.52-0.56) 0.087 (0.047-0.12) a. For each image, the question was asked: "Is image quality sufficient for classification?" Values indicate the percentage of images with adequate quality among all images in each category. Figure 1. Workflow for Progressive GAN training. A. WSIs were annotated by pathologists and small image patches of size 1024 x 1024 were extracted and augmented with random rotation and mirroring. B-C. The image patches were used as the set of real images for training a Progressive GAN model, which begins with synthesizing very small images (4 x 4 pixels) and increases the size and resolution as training progresses. D. Following GAN training, synthetic images were generated and these, along with real training images, were preprocessed (through center cropping, scaling, and color normalization) and used to train a classifier to predict cancer histotype. This article is protected by copyright. All rights reserved.
[1]
59w
All experiments were conducted in accordance with the Declaration of Helsinki and the International Ethical Guidelines for Biomedical Research Involving Human Subjects. Anonymized archival tissue samples were retrieved from the pathology archive at the BC Cancer Ovarian Care Research Program (OVCARE), University of British Columbia and Vancouver General Hospital, and were digitized after approval by the institutional ethics boards.
[2]
76w
For synthesizing histopathology images we trained a generative adversarial network on each of the histological subtype of cancers used. The images patches of 1024x1024 pixels were converted to the TFRecord format and used as the real image set in the discriminator training. The GAN training loop alternates between training the generator and the discriminator (with the other network's parameters fixed), and attempts to solve the following two-player minimax game, with the value function given by V(G,D):
[3]
211w
We used the Progressive GAN (ProGAN) training framework [11], based on the openly available code from the authors' implementation of this model All codes were implemented in Python and used the Tensorflow deep learning library [24]. Training was performed on 4 Nvidia Tesla V100 Graphical Processing Units (GPUs), and took approximately 3 days for each GAN. We did not alter the standard training parameters in the authors' code. In brief, the networks were mainly composed of 3x3 convolutional layers with the leaky ReLU activation function, with upsampling and downsampling layers used in the generator and discriminator, respectively. In the discriminator, there was also a single mini-batch standard deviation layer inserted towards the end of the network, which helps to increase the variation in training data captured by the networks. Training was performed for a total of 12,000,000 images processed, with the batch size and learning rate adjusted based on the image resolution. The Adam optimizer was used [25] with the initial learning rate=0.001 (this increased slightly at higher resolution images as per the standard training parameters), β 1 =0, β 2 =0.99, and ε=10 -8 . Following the completion of GAN training, the trained generator was used to synthesize images that were sent for pathologist review and used for data augmentation.
[4]
131w
For the pathologist review and classification of synthetic images we created two online surveys using Google Forms (https://www.google.com/forms/about/)-one for the TCGA cancer types (low grade glioma, lung SCC, hepatocellular carcinoma, papillary thyroid carcinoma, and clear cell renal cell carcinoma) and one for the ovarian carcinoma subtypes (high-grade serous, low-grade serous, endometrioid, clear cell, and mucinous). The surveys included 150 synthetic and 150 real images and were evenly split between the five histotypes. The synthetic images were generated as consecutive images using the trained generator network, while the real images were selected randomly from the training images. In order to minimize the chance of bias, the images were not screened by the survey creators. Board-certified pathologists were tasked with reviewing each of the images and answering the following three questions about each:
[5]
15w
1. What is the histotype? 2. Is the image of sufficient quality to enable classification?
[6]
110w
The sample size for the pathologist surveys was calculated using the work of Rotondi and Donner for an interobserver variability study [29]. Specifically, the survey was powered to detect a difference of 0.1 between the kappa value for histological classification of real images versus the kappa value for classification of synthetic images. Table S2 (supplementary material) shows the minimum sample size for real and synthetic images for 4 and 5 pathologists; for more details regarding choice of sample size see Supplementary methods 1. The Cohen's kappa statistic is a metric to measure the degree of agreement between two raters that accounts for the extent of agreement expected by chance alone.
[7]
73w
The suggested interpretation of kappa values (regarding level of agreement between two observers) are as follows: ≤ 0 none, 0.01-0.20 none to slight, 0.21-0.40 fair, 0.41-0.60 moderate, 0.61-0.80 substantial, This article is protected by copyright. All rights reserved. and 0.81-1.00 near perfect [30,31]. The Fleiss' kappa statistic is a generalization of Cohen's kappa that can be used for any number of raters, while Cohen's kappa can only be used for two raters [29].
[8]
9w
This article is protected by copyright. All rights reserved.
[9]
172w
We compared the performance of various generative models including Progressive GANs, variational autoencoder [33], enhanced super resolution GAN (ESRGAN) [34], and deep texture synthesis (see This article is protected by copyright. All rights reserved. Supplementary Methods 2) [35]. Upon visual inspection and quantitative comparison, we found that Progressive GANs were the only method that yielded high-resolution images (1024x1024 pixels; Figure 2 and supplementary material, Figures S1 to S3) with low Frechet Inception Distance (FID) scores (lower scores correlate to better images) and high information entropy, suggesting that high quality and visually recognizable images similar to real ones could be generated (supplementary material, Tables S3 and S4, Figure S4 and Supplementary methods 3 and 4). As such, we focused our efforts on the systematic evaluation of images generated by Progressive GANs. In addition, we experimented with using both individual Progressive GANs for each histological subtype of cancer and class-conditional GANs with multiple histological labels, and found that the GANs trained on a single subtype generated qualitatively more realistic images, with fewer artifactual distortions.
[1]
42w
For the TCGA cancer type survey, 4 pathologists (SY, DF, BS, MR) participated, from both academic and community practice settings, with 1-10 years of experience in practice. These pathologists had various areas of subspecialty expertise, including neuropathology, lung, renal, gastrointestinal, and molecular.
[2]
27w
For the ovarian cancer survey, 5 gynecologic subspecialty pathologists (BG, PI, CPH, AM, NS) participated, all from academic settings, and with 6-30 years of experience in practice.
[3]
39w
For evaluating the utility of synthetic images as data augmentation, we trained deep learning classifiers to differentiate ovarian carcinoma subtypes as well as five cancer histotypes from TCGA. Each image This article is protected by copyright. All rights reserved.
[4]
27w
classifier was trained and tested on an additional set of 52 (ovarian) and 108 (TCGA) WSIs that were not included in the training set of the GANs.
[5]
84w
Image patches were pre-processed as follows: ovarian cancer images were centre-cropped to 768 x 768 pixels and down-sampled to 256 x 256, and TCGA images were down-sampled to 512 x 512. We then applied HSV colour jitter (during training only) with parameters adopted from [26] (brightness with a maximum delta of 0.25, contrast with a maximum delta of 0.75, saturation with a maximum delta of 0.25, and hue with a maximum delta of 0.04), and normalized to RGB pixel values between -1 and 1.
[6]
248w
Using these images we randomly created ten splits of equal numbers of patients allocated to training/validation and testing data sets. The baseline training image set (average n = 1848 image patches for ovarian and n = 1967 for TCGA) was then augmented with either real images from the GAN training set or synthetic images generated by the GANs (n = 40,000 patches for ovarian and n = 2,000 for TCGA). Therefore, in total three different classifiers with identical hyperparameter settings were trained on each of the 10 patient splits, one for each of the following: baseline train images only (referred to as baseline setx), baseline images augmented with real images (referred to as baseline + real set), and baseline images augmented with synthetic images (referred to as baseline + synthetic set). On average, the validation and testing sets contained 7,850 and 18,802 image patches for ovarian and 14,617 and 12,994 image patches for TCGA, respectively. The VGG19 network [27] with batch normalization was chosen for the classifier based on its ease of use and robust performance. The PyTorch [28] implementation was used with a modified last fully connected layer to classify the five classes (either five subtypes of ovarian carcinoma or five TCGA histotypes), with softmax activation to obtain the categorical distribution. We initialized each VGG19 model with weights trained on ImageNet, and all the weights in the model were optimized during back propagation. The Adam optimizer [25] was used with learning rate=2*10 -4 , beta1=0.9, and beta2=0.999.
[7]
32w
Training was done with batch size=64 for 10 epochs and, using a single Tesla V100 GPU, took approximately 2 and 5 hours per model for the augmented TCGA and OVCARE datasets, respectively.
[8]
54w
Classification results are reported on the testing set using the weights of the VGG19 model taken from the training epoch with the highest validation accuracy. As the purpose was not to achieve the greatest absolute performance, but rather to evaluate the utility of synthetic data for improving performance, no significant hyperparameter tuning was performed.