基于快速地标采样的大规模谱聚类算法

发布时间：2018-05-01 07:29

本文选题：地标点采样 + 大数据　；参考：《电子与信息学报》2017年02期

【摘要】：为避免传统谱聚类算法高复杂度的应用局限,基于地标表示的谱聚类算法利用地标点与数据集各点间的相似度矩阵,有效降低了谱嵌入的计算复杂度。在大数据集情况下,现有的随机抽取地标点的方法会影响聚类结果的稳定性,k均值中心点方法面临收敛时间未知、反复读取数据的问题。该文将近似奇异值分解应用于基于地标点的谱聚类,设计了一种快速地标点采样算法。该算法利用由近似奇异向量矩阵行向量的长度计算的抽样概率来进行抽样,同随机抽样策略相比,保证了聚类结果的稳定性和精度,同k均值中心点策略相比降低了算法复杂度。同时从理论上分析了抽样结果对原始数据的信息保持性,并对算法的性能进行了实验验证。
[Abstract]:In order to avoid the high complexity of the traditional spectral clustering algorithm, the spectral clustering algorithm based on landmarks can effectively reduce the computational complexity of spectral embedding by using the similarity matrix between the ground punctuation points and the data sets. In the case of big data set, the existing random sampling ground punctuation methods will affect the stability of the clustering results and the k-means centroid method will face the problem of the unknown convergence time and the problem of repeatedly reading the data. In this paper, the approximate singular value decomposition is applied to the spectral clustering based on geopunctuation, and a fast punctuation sampling algorithm is designed. The algorithm uses the sampling probability calculated by the length of the approximate singular vector matrix row vector to carry out the sampling. Compared with the random sampling strategy, the stability and accuracy of the clustering results are guaranteed. Compared with the k-means center point strategy, the algorithm complexity is reduced. At the same time, the information retention of the sampling results to the original data is analyzed theoretically, and the performance of the algorithm is verified experimentally.
【作者单位】：解放军信息工程大学;数学工程与先进计算国家重点实验室;
【基金】：国家973计划(2012CB315905) 国家自然科学基金(61502527,61379150)~~
【分类号】：TP311.13

【相似文献】