首页 | 官方网站   微博 | 高级检索  
     

基于可变网格划分的密度偏差抽样算法
引用本文:盛开元,钱雪忠,吴秦.基于可变网格划分的密度偏差抽样算法[J].计算机应用,2013,33(9):2419-2422.
作者姓名:盛开元  钱雪忠  吴秦
作者单位:江南大学 物联网工程学院,江苏 无锡 214122
基金项目:国家自然科学基金资助项目,江苏省科技支撑计划项目
摘    要:简单随机抽样是在分析处理大规模数据集时最常用的数据约简方法,但该方法在处理内部分布不均匀的数据集时容易造成类的丢失。基于固定网格划分的密度偏差抽样算法虽能有效解决该问题,但其速度及效果易受网格划分粒度影响。为此提出了基于可变网格划分的密度偏差抽样算法,根据原始数据集每一维的分布特征确定该维相应的划分粒度,进而构建与原始数据集分布特征一致的网格空间。实验结果表明,在可变网格划分的基础上进行密度偏差抽样,样本质量明显提升,而且相对于基于固定网格划分的密度偏差抽样算法,抽样效率亦有所提高。

关 键 词:密度偏差抽样    可变网格划分    数据挖掘    大规模数据集    聚类
收稿时间:2013-04-08
修稿时间:2013-04-29

Density biased sampling algorithm based on variable grid division
SHENG Kaiyuan , QIAN Xuezhong , WU Qin.Density biased sampling algorithm based on variable grid division[J].journal of Computer Applications,2013,33(9):2419-2422.
Authors:SHENG Kaiyuan  QIAN Xuezhong  WU Qin
Affiliation:School of Internet of Things Engineering, Jiangnan University, Wuxi Jiangsu 214122, China
Abstract:As the most commonly used method of reducing large-scale datasets, simple random sampling usually causes the loss of some clusters when dealing with unevenly distributed dataset. A density biased sampling algorithm based on grid can solve these defects, but both the efficiency and effect of sampling can be affected by the granularity of grid division. To overcome the shortcoming, a density biased sampling algorithm based on variable grid division was proposed. Every dimension of original dataset was divided according to the corresponding distribution, and the structure of the constructed grid was matched with the distribution of original dataset. The experimental results show that density biased sampling based on variable grid division can achieve higher quality of sample dataset and uses less execution time of sampling compared with the density biased sampling algorithm based on fixed grid division.
Keywords:density biased sampling  variable grid division  data mining  large-scale dataset  clustering
本文献已被 万方数据 等数据库收录!
点击此处可从《计算机应用》浏览原始摘要信息
点击此处可从《计算机应用》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司    京ICP备09084417号-23

京公网安备 11010802026262号