电子学报 ›› 2009, Vol. 37 ›› Issue (2): 278-284.

• 论文 • 上一篇    下一篇

基于主题分析的文本分割技术研究

刘 铭, 王晓龙, 刘远超   

  1. 哈尔滨工业大学计算机科学与技术学院,黑龙江哈尔滨 150001
  • 收稿日期:2007-12-10 修回日期:2008-09-16 出版日期:2009-02-25 发布日期:2009-02-25

Research on Text Segmentation Based on Topic Analysis

LIU Ming, WANG Xiao-long, LIU Yuan-chao   

  1. School of Computer Science and Technology,Harbin Institute of Technology,Harbin,Heilongjiang 150001,China
  • Received:2007-12-10 Revised:2008-09-16 Online:2009-02-25 Published:2009-02-25

摘要: 本文提出一种新颖的文本分割算法,算法首先将待分割文档划分为若干片段的集合,然后构造全文词汇链分析文中描述的多个子主题,并通过构造片段对子主题的覆盖图将描述相同子主题的相似片段归类.针对段落分割点可能落在片段内部的情况,算法对片段进行二次划分.实验表明:在对文档进行主题分析后,算法能够过滤掉与主题无关的特征对分割结果的干扰;构造的片段对子主题的覆盖图融合了相邻及相间片段的相似性,加大了划分的准确度;对片段进行二次划分使得分割的结果更加合理.

关键词: 主题分析, 词汇链, 知网, 二次划分

Abstract: A novel topic segmentation algorithm is proposed in this paper.This algorithm first partitions text into some blocks.After that it constructs whole-length lexical chains to analyze multiple subtopics of this text.By constructing graph which describes blocks covering subtopics,the similar blocks which describe same subtopic can be classified.In order to solve the situations that segmentation points drop inside blocks,it segments blocks again.Experiment results demonstrate that by analyzing topic of text,this algorithm can remove interferences,which are aroused by irrelative features,from segmentation results.By constructing graph which describes blocks covering subtopics,it can mix similarities of adjacent and disconnected blocks together,and increases segmentation precision.The second segmentation makes segmentation results more reasonable.

Key words: topic analysis, lexical chain, HowNet, second segmentation

中图分类号: