<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE ArticleSet PUBLIC "-//NLM//DTD PubMed 2.7//EN" "https://dtd.nlm.nih.gov/ncbi/pubmed/in/PubMed.dtd">
<ArticleSet>
<Article>
<Journal>
				<PublisherName>Shahrood University of Technology</PublisherName>
				<JournalTitle>Journal of AI and Data Mining</JournalTitle>
				<Issn>2322-5211</Issn>
				<Volume>6</Volume>
				<Issue>2</Issue>
				<PubDate PubStatus="epublish">
					<Year>2018</Year>
					<Month>07</Month>
					<Day>01</Day>
				</PubDate>
			</Journal>
<ArticleTitle>Extracting Predictor Variables to Construct Breast Cancer Survivability Model with Class Imbalance Problem</ArticleTitle>
<VernacularTitle></VernacularTitle>
			<FirstPage>263</FirstPage>
			<LastPage>276</LastPage>
			<ELocationID EIdType="pii">1058</ELocationID>
			
<ELocationID EIdType="doi">10.22044/jadm.2017.5061.1609</ELocationID>
			
			<Language>EN</Language>
<AuthorList>
<Author>
					<FirstName>S.</FirstName>
					<LastName>Miri Rostami</LastName>
<Affiliation>Faculty of computer and IT Engineering, Shiraz University of Technology, Shiraz, Iran.</Affiliation>

</Author>
<Author>
					<FirstName>M.</FirstName>
					<LastName>Ahmadzadeh</LastName>
<Affiliation>Faculty of computer and IT Engineering, Shiraz University of Technology, Shiraz, Iran.</Affiliation>

</Author>
</AuthorList>
				<PublicationType>Journal Article</PublicationType>
			<History>
				<PubDate PubStatus="received">
					<Year>2016</Year>
					<Month>11</Month>
					<Day>19</Day>
				</PubDate>
			</History>
		<Abstract>Application of data mining methods as a decision support system has a great benefit to predict survival of new patients. It also has a great potential for health researchers to investigate the relationship between risk factors and cancer survival. But due to the imbalanced nature of datasets associated with breast cancer survival, the accuracy of survival prognosis models is a challenging issue for researchers. This study aims to develop a predictive model for 5-year survivability of breast cancer patients and discover relationships between certain predictive variables and survival. The dataset was obtained from SEER database. First, the effectiveness of two synthetic oversampling methods Borderline SMOTE and Density based Synthetic Oversampling method (DSO) is investigated to solve the class imbalance problem. Then a combination of particle swarm optimization (PSO) and Correlation-based feature selection (CFS) is used to identify most important predictive variables. Finally, in order to build a predictive model three classifiers decision tree (C4.5), Bayesian Network, and Logistic Regression are applied to the cleaned dataset. Some assessment metrics such as accuracy, sensitivity, specificity, and G-mean are used to evaluate the performance of the proposed hybrid approach. Also, the area under ROC curve (AUC) is used to evaluate performance of feature selection method. Results show that among all combinations, DSO + PSO_CFS + C4.5 presents the best efficiency in criteria of accuracy, sensitivity, G-mean and AUC with values of 94.33%, 0.930, 0.939 and 0.939, respectively.</Abstract>
		<ObjectList>
			<Object Type="keyword">
			<Param Name="value">Breast Cancer</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">survival</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Class Imbalance Problem</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">oversampling technique</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Feature Selection</Param>
			</Object>
		</ObjectList>
<ArchiveCopySource DocType="pdf">https://jad.shahroodut.ac.ir/article_1058_2647c65fe9ab0e31072d03a6fb22fdc2.pdf</ArchiveCopySource>
</Article>
</ArticleSet>
