1;3409;0c A generalised regression algorithm for Web page categorisation

A generalised regression algorithm for Web page categorisation

Neural Computing & Applications, vol. 13, no. 3, 2004
Pages: 229-236DOI: 10.1007/s00521-004-0409-0

NCA

bibtex

This paper proposes an information system that classifies Web pages according a taxonomy, which is mainly used from seven search engines/directories. The proposed classifier is a four-layer generalised regression neural network (GRNN) that aims to perform the information segmentation according to information filtering techniques using content descriptor vectors. Eight categories of Web pages were used in order to evaluate the robustness of the method, while no restrictions were imposed except for the language of the content, which is English. The system can be used as an assistant and consultative tool for classification purposes as well as for estimating the population of Web pages at any given point in time.