欢迎访问林草资源研究
技术应用

森林资源抽样调查缺失数据填充方法

  • 刘菲 ,
  • 李明阳 ,
  • 刘雅楠 ,
  • 江一帆 ,
  • 王子
展开
  • 南京林业大学 林学院,南京 210037
刘菲(1994-),女,安徽滁州人,在读硕士,从事3S技术应用方面的研究。Email: 121126082@qq.com

收稿日期: 2018-09-17

  修回日期: 2018-12-10

  网络出版日期: 2020-09-27

基金资助

国家自然科学基金项目“基于情景分析与多目标决策的南方集体林长期经营规划方法研究”(31770679)

Filling Method for Missing Data of Forest Resource Sampling Investigation

  • Fei LIU ,
  • Mingyang LI ,
  • Yanan LIU ,
  • Yifan JIANG ,
  • Zi WANG
Expand
  • College of Forestry,Nanjing Forestry University,Nanjing,Jiangsu 210037,China

Received date: 2018-09-17

  Revised date: 2018-12-10

  Online published: 2020-09-27

摘要

在森林资源抽样调查中数据缺失现象时常发生,为了提高数据分析的准确性,有必要对缺失数据填充方法进行研究。以浙江省临安市1996年Landsat-5 TM影像及同期县级森林资源连续监测固定样地数据为主要信息源,以样地内林木平均胸径为缺失因子,在对其空间自相关分析的基础上,采用十折交叉验证法对缺失数据进行空间、非空间和基于遥感估测模型填充以及精度评价。结果表明:1)研究区样地林木平均胸径的Moran’s I系数为0.21,空间分布表现出较强的空间自相关性;2)遥感估测模型中K-近邻算法的填充精度最高,其次为随机森林、空间填充的克里金内插,非空间的期望极大化算法填充精度最低;3)克里金内插的4个半方差理论模型中,球状模型填充精度最高,相关系数(0.632 5)最高,平均绝对误差(2.049 3cm)和均方根误差(3.809 3cm)最低;4)按照填充精度由高到低的顺序,4种性能较好的数据填充方法依次为:K-近邻算法>随机森林>克里金内插>距离权重反比。在地势形态复杂、海拔差异较大的临安境内,K-近邻算法较适合样地林木平均胸径因子的缺失数据填充。

本文引用格式

刘菲 , 李明阳 , 刘雅楠 , 江一帆 , 王子 . 森林资源抽样调查缺失数据填充方法[J]. 林草资源研究, 2018 , 0(6) : 130 -137 . DOI: 10.13466/j.cnki.lyzygl.2018.06.021

Abstract

The phenomenon of data loss often occurs in forest resource sampling investigation.So it is necessary to study the filling method of missing data in order to improve the accuracy of the data analysis.Linan County located in Zhejiang Province was chosen as the case study area.Landsat-5 TM image in 1996 and County-level fixed plot data of forest resources continuous detection in the same period were used as the main information,and the average DBH(Diameter at Breast Height) of trees in sample plot as the missing factor to make spatial filling,non-spatial filling,model filling of remote sensing estimation for missing data.And 10 fold cross-validation method on the basis of spatial autocorrelation analysis of the average DBH of trees in sample plot was employed to make accuracy evaluation.The results show that:(1) The Moran’I coefficient of the average DBH of sample plot trees in study area is 0.21 and its spatial distribution shows strong spatial autocorrelation;(2)The filling accuracy of K-Nearest Neighbor of remote sensing estimation models is the highest,the second is Random Forest followed by the Kriging Interpolation of spatial filling.However,the filling accuracy of expectation maximization algorithm of non-spatial fillings is the lowest;(3)Among four semi-variance models of Kriging interpolation,the filling accuracy of spherical model is higher than any other models.Its correlation coefficient constitutes 0.632 5,the mean absolute error makes up 2.049 3 centimeters and the root mean square error accounts for 3.809 3 centimeters;(4)According to the order of filling accuracy from high to low,four priority filling methods of missing data includes:K-Nearest Neighbor,Random Forest,Kriging Interpolation and Inverse Distance Weighting.It is the K-Nearest Neighbor that is most suitable for filling missing data of the average DBH of sample plot trees in Linan with complex topography and great different altitudes.

参考文献

[1] 李明阳, 刘敏, 刘米兰. 基于GIS的森林调查因子地统计学分析[J]. 南京林业大学学报:自然科学版, 2010,34(6):66-70.
[2] Dempster A P, Laird N M, Rubin D B. Maximan likelihood estimation from incomplete data via the algorithm[J]. Journal of the Royal Statistical Society Series B-statistical Methodology, 1977,39:1-38.
[3] Rubin D B. Multiple imputation after 18+ years[J]. Journal of the American Statistical Association, 1996,91(434):473-489.
[4] Rubin D B. Multiple imputation a primer[J]. Statistical Methods in Medical Research, 1999,8(1):3-15.
[5] Chiu H Y, Sedransk J. A Bayesian procedure for imputing missing values in sample surveys[J]. Journal of the American Statistical Association, 1986,81(395):667-676.
[6] Astebro T, Chen G. How to deal with missing categorical data:Test of a simple Bayesian method[J]. Organizational Research Methods, 2010,6(3):309-327.
[7] Tara B, Matti M. Missing data in forest ecology and management:Advances in quantitative methods[J]. Forest Ecology and Management, 2012,271:1-2.
[8] 何红艳, 郭志华, 肖文发. 降水空间插值技术的研究进展[J]. 生态学杂志, 2005,24(10):1187-1191.
[9] 靳国栋, 刘衍聪, 牛文杰. 距离权重反比插值法和克里金插值法的比较[J]. 长春工业大学学报, 2003,24(3):53-57.
[10] 张连强, 赵有中, 欧阳宗继, 等. 运用地理因子推算山区局地降水量的研究[J]. 中国农业气象, 1996,17(2):6-10.
[11] 王丹丹. 空间统计分析及其在农用地分等中的应用[D]. 西安:长安大学, 2008.
[12] 张文彤, 董伟. SPSS统计分析高级教程[M]. 北京: 高等教育出版社, 2004.
[13] 蒋云姣, 胡曼, 李明阳, 等. 县域尺度森林地上生物量遥感估测方法研究[J]. 西南林业大学学报, 2015,35(6):53-59.
[14] 荣媛, 刘任琪, 李明阳, 等. 基于星载高光谱数据的南京新济州湿地土壤有机质估测研[J].西南林业大学学报, 2017(6):171-177.
[15] Breiman L. Random forests[J]. Machine Learning, 2001,45(1):5-32.
[16] Goldstein B A, Hubbard A E, Cutle A, et al. An application of random forests to genome-wide association dataset:methodological considerations & new findings[J]. BMC Genetics, 2010,11(1):49-61.
[17] Fullerr R M, Devereux B J, Gillings S, et al. Indices of bird-habitat preference from field surveys of birds and remote sensing of land cover:a study of south-eastern England with wider implications for conservation and biodiversity assessment[J]. Global Ecology & Biogeography, 2005,14(3):223-229.
[18] 李明阳, 余超, 张密芳, 等. 紫金山风景林生物量及驱动因素时间轨迹分析[J]. 北京林业大学学报, 2015,37(2):1-7.
[19] David E. Statistics in Geography[M]. Oxford:Oxford Basil Blackwell Ltd, 1985.
[20] 戴前石, 刘金山. 青藏高原贡觉县森林规划设计因子的地统计学分析[J]. 西南林业大学学报, 2017,37(3):146-151.
[21] 陈伟强, 刘国顺, 华一新, 等. 基于GIS的河南省典型烟区土壤养分时空变异分析[J]. 河南农业科学, 2007,36(11):70-75.
文章导航

/