Study on Outliers Influence in NIR Quantitative Analysis Model
ZHENG Feng1, LIU Li-ying1, LIU Xiao-xi2, LI Ye1, SHI Xiao-guang1, ZHANG Guo-yu1, HUAN Ke-wei1*
1. Changchun University of Science and Technology, Changchun 130022, China 2. Institute of Scientific and Technical Information in Jilin Province, Changchun 130000, China
Abstract:As a secondary analysis method, reproducibility and reliability of near-infrared spectroscopy (NIRS) quantitative analysis are quite dependent on modelling process. In this paper,it is focused on outlier analysis for protein quantitative model of wheat based on NIRS. The purpose is to discuss the outlier effect in modelling process of complex sample set. The indicator of outliers is the deviation between two interpretative percentage curves in partial least squares regression (PLSR) modelling, when two percentage curves have significant deviation or departure point, the sample set should include the outliers. The innovative research work is the analysis and treatment of outliers. On the basis of sub-model ergodic calculation method, outliers can be gradually identified and picked-up. The standard deviation of model’s prediction residual is used as the reference graduation to distinguish the degree of deviation. According to the degree of deviation from sample population, outliers can also be divided into significant outliers, relative outliers and potential outliers. In this paper, the significant outliers of the sample set are about 7.8%, and the relative outliers are about 15.6%. The outliers will pull normal samples apart from the ideal fitting line and make the dispersity increase. No matter modelling with removed outliers or weighted samples, the purpose is to make the fitting results of quantitative analysis modelling more inclined to majority samples, while reducing or eliminating the impact of outliers.
Key words:Near infrared spectroscopy;Outlier analysis;Gray system;Sub-model population learning
郑 峰1,刘丽莹1,刘小溪2,李 野1,石晓光1,张国玉1,宦克为1* . 近红外光谱定量分析模型的样本影响研究 [J]. 光谱学与光谱分析, 2016, 36(11): 3523-3529.
ZHENG Feng1, LIU Li-ying1, LIU Xiao-xi2, LI Ye1, SHI Xiao-guang1, ZHANG Guo-yu1, HUAN Ke-wei1* . Study on Outliers Influence in NIR Quantitative Analysis Model . SPECTROSCOPY AND SPECTRAL ANALYSIS, 2016, 36(11): 3523-3529.
[1] CHU Xiao-li, LU Wan-zhen(褚小立, 陆婉珍). Spectroscopy and Spectral Analysis(光谱学与光谱分析), 2014, 34(10):2595. [2] HAO Yong, CAI Wen-sheng, SHAO Xue-guang(郝 勇, 蔡文生, 邵学广), Chemical Journal of Chinese Universities(高等学校化学学报), 2009, 30:28. [3] LIANG Yi-zeng, XU Qing-song(梁逸曾, 徐青松). Instrumental Analysis of Complex Systems ——White, Gray and Black Analytical Systems and Their Multivariate Methods(复杂体系仪器分析——白、灰、黑分析体系及其多变量解析方法). Beijing: Chemical Industry Press(北京:化学工业出版社), 2012. [4] LI Hongdong, Liang Yizeng, Cao Dongsheng, et al. Trac Trends in Analytical Chemistry, 2012, 38(9): 154. [5] Vladimir N Vapnik. Statistical Learning Theory. New York: Wiley-Interscience, 1998. [6] Klaus Danzer. Analytical Chemistry: Theoretical and Metrological Fundamentals. New York:Springer-Verlag Berlin Heidelberg Press, 2007. [7] Tomaso Poggio, Ryan Rifkin, Sayan Mukherjee, et al. Nature,2004, 428:419. [8] Deng Baichuan, Yun Yonghuan, Liang Yizeng. Chemometrics and Intelligent Laboratory Systems, 2015, 149: 166. [9] Beckman R J,Cook R D. Technometrics, 1983, 25(2): 119. [10] BAI Wen-liang, ZHANG Jun, GAN Feng, et al. Computers and Applied Chemistry(计算机与应用化学),2010, 27(11):1476.