基于《知网》的词语相似度算法研究 Study on HowNet-Based Word Similarity Algorithm期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

基于《知网》的词语相似度算法研究

引用本文：	刘青磊,顾小丰.基于《知网》的词语相似度算法研究[J].中文信息学报,2010,24(6):31-37.

作者姓名：	刘青磊顾小丰

作者单位：	电子科技大学计算机科学与工程学院,四川成都 611731

基金项目：	国家863计划资助项目，国家自然科学基金资助项目，四川省科技厅资助项目

摘要：	基于《知网》的词语(句子)相似度计算通常是把义原(词语)之间的最优匹配做为运算的基本单位的,最终的整体相似度数值可由每一部分的相似度值通过适当的加权计算合成而来,这样的做法往往会造成一些匹配对的信息重复和结构不合理。针对这个问题,该文通过统计出两个直接义原集合间的共有信息(共性)和差异信息(个性)来计算集合的相似度,并把此方法引入到词语(句子)的相似度计算中去。最终的实验比对结果表明该文所采用的方法更为稳定和有效。
关键词：	《知网》词语相似度句子相似度共有信息差异信息
Study on HowNet-Based Word Similarity Algorithm

LIU Qinglei,GU Xiaofeng.Study on HowNet-Based Word Similarity Algorithm[J].Journal of Chinese Information Processing,2010,24(6):31-37.

Authors:	LIU Qinglei GU Xiaofeng

Affiliation:	Department of Computer Science and Technology, University of Electronic Science and Technology of China, Chengdu, Sichuan 611731, China

Abstract:	Word (sentence) similarity computing based on the “HowNet” usually treats the optimal matches between the primitives or words as the basic unit, and the ultimate outcome can be the sum of weighted counts. However, this approach often results in the information duplication and some irrational constructions. To deal with these issues, this paper propose to calculate the similarity of sets by the statistics on common information (commonality) and the different information (differences) between the two sets of direct primitives. Moreover, the paper introduces this measure into the calculation of sentence similarity. The final experimental analysis shows that the proposed method is more stable and effective. Key wordsHowNet; word similarity; sentence similarity; common information; different information

Keywords:	HowNet word similarity sentence similarity common information different information
本文献已被万方数据等数据库收录！
	点击此处可从《中文信息学报》浏览原始摘要信息
	点击此处可从《中文信息学报》下载全文

设为首页 | 免责声明 | 关于勤云 | 加入收藏