APP下载

基于Delphi的Web文本获取方法

2016-03-21刘建培

计算机时代 2016年3期

刘建培

摘 要: 提出基于delphi的Web文本获取方法,从网页中获取Web页面格式的源文件(.html文件),分析它的结构信息,处理它的控制符,通过分析过滤源文件的格式来提取网页中的文本信息。利用标点符号对文本信息进行章节、段落、句子等预处理,将文本信息转换成句子序列,让用户快速地定位到需要了解的内容,从而让用户远离钓鱼网站、恶意广告、欺诈信息以及在浏览网页内容时产生的骚扰,提高互联网体验。

关键词: Delphi; 文本获取; HTML; 控制符

中图分类号:TP391 文献标志码:A 文章编号:1006-8228(2016)03- -03

A Web text acquisition method

Liu Jianpei

(Educational technology center of Guangdong university of finance & economics, Guangzhou, Guangdong 510320, China)

Abstract: In this paper, a method of Web text acquisition with Delphi is proposed, which obtains the source files of the Web page format (.Html file) from the Web page, analyzes its structure information, deals with its control character, and extracts the text information from the Web page by analyzing and filtering the source files formats. The method makes use of punctuation marks to preprocess the text information for sections, paragraphs and sentences, converts the text information into sentence sequences, which allows the users to quickly navigate to the contents needed to know, allows the users to stay away from phishing sites, malicious advertising, fraud information and the harassment generated by browsing the content of Web pages, and improves their Internet experience.

Key words: Delphi; text acquisition; HTML; control character

0 引言

互联网时代,各式各样的站点中积累了丰富的文档资料,其中不仅有名目繁多的技术资料和新闻资讯,还有众多用户的观点和评论。人们浏览网页文档资料获得所需要的信息,也难免受到钓鱼网站、恶意广告、欺诈信息及各种骚扰,用户为个人隐私及数据安全而烦恼。本文提出基于delphi的Web文本获取,快速地定位需要了解的内容,从而让用户远离烦恼,提高互联网体验。

1 实现步骤

⑴ 获取论坛文档:输入一个论坛文档的网址,获取网页源码,对网页源码过滤,最终获取文档文本。

⑵ 文本处理:能利用标点符号对文档进行章节、段落、句子等预处理工作,将文档转换成句子序列。

2 获取Web文本

系统首先在线从网页中获取Web页面[4]格式的源文件,通过分析过滤源文件(.html文件)的格式,提取网页中的文本信息。

网页信息是用HTML(Hypertext Markup Language)语言书写的,我们要对其中的文本信息进行提取,必须首先分析它的结构信息[5]。……

登录APP查看全文