1;3409;0c Automating the Construction of Internet Portals with Machine Learning

Automating the Construction of Internet Portals with Machine Learning

Information Retrieval, vol. 3, no. 2, 2000
Pages: 127-163

IR

bibtex

. Domain-specic internet portals are growing in popularity because they gather content from the Web and organize it for easy access, retrieval and search. For example, www.campsearch.com allows complex queries by age, location, cost and specialty over summer camps. This functionality is not possible with general, Web-wide search engines. Unfortunately these portals are dicult and time-consuming to maintain. This paper advocates the use of machine learning techniques to greatly automate the creation and maintenance of domain-specic Internet portals. We describe new research in reinforcement learning, information extraction and text classication that enables ecient spidering, the identication of informative text segments, and the population of topic hierarchies. Using these techniques, we have built a demonstration system: a portal for computer science research papers. It already contains over 50,000 papers and is publicly available at www.cora.justresearch.com. These techniques are ...