首页 | 官方网站   微博 | 高级检索  
     

基于《知网》的词语相似度算法研究
引用本文:刘青磊,顾小丰.基于《知网》的词语相似度算法研究[J].中文信息学报,2010,24(6):31-37.
作者姓名:刘青磊  顾小丰
作者单位:电子科技大学 计算机科学与工程学院,四川 成都 611731
基金项目:国家863计划资助项目,国家自然科学基金资助项目,四川省科技厅资助项目
摘    要:基于《知网》的词语(句子)相似度计算通常是把义原(词语)之间的最优匹配做为运算的基本单位的,最终的整体相似度数值可由每一部分的相似度值通过适当的加权计算合成而来,这样的做法往往会造成一些匹配对的信息重复和结构不合理。针对这个问题,该文通过统计出两个直接义原集合间的共有信息(共性)和差异信息(个性)来计算集合的相似度,并把此方法引入到词语(句子)的相似度计算中去。最终的实验比对结果表明该文所采用的方法更为稳定和有效。

关 键 词:《知网》  词语相似度  句子相似度  共有信息  差异信息  

Study on HowNet-Based Word Similarity Algorithm
LIU Qinglei,GU Xiaofeng.Study on HowNet-Based Word Similarity Algorithm[J].Journal of Chinese Information Processing,2010,24(6):31-37.
Authors:LIU Qinglei  GU Xiaofeng
Affiliation:Department of Computer Science and Technology, University of Electronic Science
and Technology of China, Chengdu, Sichuan 611731, China
Abstract:Word (sentence) similarity computing based on the “HowNet” usually treats the optimal matches between the primitives or words as the basic unit, and the ultimate outcome can be the sum of weighted counts. However, this approach often results in the information duplication and some irrational constructions. To deal with these issues, this paper propose to calculate the similarity of sets by the statistics on common information (commonality) and the different information (differences) between the two sets of direct primitives. Moreover, the paper introduces this measure into the calculation of sentence similarity. The final experimental analysis shows that the proposed method is more stable and effective.
Key wordsHowNet; word similarity; sentence similarity; common information; different information
Keywords:HowNet  word similarity  sentence similarity  common information  different information
 
        
 
        
 
        
本文献已被 万方数据 等数据库收录!
点击此处可从《中文信息学报》浏览原始摘要信息
点击此处可从《中文信息学报》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司    京ICP备09084417号-23

京公网安备 11010802026262号