首页 | 本学科首页   官方微博 | 高级检索  
     检索      

中文农业搜索引擎字符编码识别
引用本文:吴乃宁,张太红,白涛.中文农业搜索引擎字符编码识别[J].新疆农业大学学报,2014(5):420-423.
作者姓名:吴乃宁  张太红  白涛
作者单位:新疆农业大学 计算机与信息工程学院,乌鲁木齐,830052
基金项目:新疆维吾尔自治区科技攻关项目
摘    要:针对农业网页中汉字编码标识混乱的情况,提出了一种综合运用编码规则和网页文本特征的字符编码识别模型。利用卡方检验算法,结合最小二乘多元线性回归方法,得到了基于网页文本特征的字符识别模型。实验结果显示,在适当的选取阈值(r =1,阈值=属于某一编码的字符数/网页总字符数)和文本特征数(≥65)的基础上,模型准确率达到100%,且结果稳定。

关 键 词:编码识别  卡方检验  多元线性回归  GB2312  Big5

Character Encoding Identification of Chinese Agriculture Search Engine
WU Nai-ning,ZHANG Tai-hong,BAI Tao.Character Encoding Identification of Chinese Agriculture Search Engine[J].Journal of Xinjiang Agricultural University,2014(5):420-423.
Authors:WU Nai-ning  ZHANG Tai-hong  BAI Tao
Institution:(College of Computer & Information Engineering, Xiniiang Agricultural University, Urumqi 830052, China)
Abstract:The character encoding identification model comprehensively using encoding rules and Web page text was put forward in accordance with the confused conditions of character encoding identification in Chinese agriculture web page.Using the chi-square test algorithm,combining with the method of least square multivariale linear regression,the model of character identification based on Web page text feature was obtained.The experimental results showed that the accuracy of the model reached 100% and the result was stable on the basis of the appropriate threshold selected (r =1,threshold=the number of characters belonging to a coding/the total number of web page)and the text feature number (≥65).
Keywords:coded identification  Chi-square test  multiple linear regression  GB2312  Big5
本文献已被 维普 万方数据 等数据库收录!
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号