基于文档发散度的作文跑题检测
发布时间:2018-03-01 07:50
本文关键词: 跑题检测 文档发散度 文本相似度 出处:《中文信息学报》2017年01期 论文类型:期刊论文
【摘要】:作文跑题检测是作文自动评分系统的重要模块。传统的作文跑题检测一般计算文章内容相关性作为得分,并将其与某一固定阈值进行对比,从而判断文章是否跑题。但是实际上文章得分高低与题目有直接关系,发散性题目和非发散性题目的文章得分有明显差异,所以很难用一个固定阈值来判断所有文章。该文提出一种作文跑题检测方法,基于文档发散度的作文跑题检测方法。该方法的创新之处在于研究文章集合发散度的概念,建立发散度与跑题阈值的关系模型,对于不同的题目动态选取不同的跑题阈值。该文构建了一套跑题检测系统,并在一个真实的数据集中进行测试。实验结果表明基于文档发散度的作文跑题检测系统能有效识别跑题作文。
[Abstract]:Composition detection is an important module of automatic scoring system. The composition of the composition detection of traditional calculation content as the correlation score, and compares it with a fixed threshold, so as to judge whether the article subject. But there is a direct relationship between the score and title actually, have obvious differences of divergent topics and non divergent topics the score, it is difficult to use a fixed threshold to determine all of the article. This paper proposes a detection method of composition, composition detection method based on document divergence. The innovation of this method is the concept of collection of divergence, divergence and relationship model is established for dynamic threshold point, topic but the different selection of different threshold. This paper constructs a set point detection system, and in a real data set for testing. The experimental results table Composition detection system in BenQ document divergence can effectively identify the topic composition.
【作者单位】: 苏州大学计算机科学与技术学院;软件新技术与产业化协同创新中心;
【基金】:国家自然科学基金(61572338)
【分类号】:H15;TP391.1
,
本文编号:1551059
本文链接:https://www.wllwen.com/wenyilunwen/yuyanyishu/1551059.html