Dependency Analysis of Accuracy Estimates in k-Fold Cross Validation

Tzu Tsung Wong, Nai Yu Yang

研究成果: Article同行評審

124 引文 斯高帕斯(Scopus)

摘要

A standard procedure for evaluating the performance of classification algorithms is k-fold cross validation. Since the training sets for any pair of iterations in k-fold cross validation are overlapping when the number of folds is larger than two, the resulting accuracy estimates are considered to be dependent. In this paper, the overlapping of training sets is shown to be irrelevant in determining whether two fold accuracies are dependent or not. Then a statistical method is proposed to test the appropriateness of assuming independence for the accuracy estimates in k-fold cross validation. This method is applied on 20 data sets, and the experimental results suggest that it is generally appropriate to assume that the fold accuracies are independent. The cross validation of non-overlapping training sets can make fold accuracies to be dependent. However, this dependence almost has no impact on estimating the sample variance of fold accuracies, and hence they can generally be assumed to be independent.

原文English
文章編號8012491
頁(從 - 到)2417-2427
頁數11
期刊IEEE Transactions on Knowledge and Data Engineering
29
發行號11
DOIs
出版狀態Published - 2017 11月 1

All Science Journal Classification (ASJC) codes

  • 資訊系統
  • 電腦科學應用
  • 計算機理論與數學

指紋

深入研究「Dependency Analysis of Accuracy Estimates in k-Fold Cross Validation」主題。共同形成了獨特的指紋。

引用此