Algoritmo rastreador web especialista nuclear

Reis, Thiago

doi:10.11606/D.85.2013.tde-07012014-134548

Home

Facilities

Master's Dissertation

DOI

https://doi.org/10.11606/D.85.2013.tde-07012014-134548

Document

Master's Dissertation

Author

Reis, Thiago (Catálogo USP)

Full name

Thiago Reis

E-mail

Institute/School/College

Instituto de Pesquisas Energéticas e Nucleares

Knowledge Area

Nuclear Technology - Reactors

Date of Defense

2013-11-12

Published

São Paulo, 2013

Supervisor

Barroso, Antonio Carlos de Oliveira (Catálogo USP)

Committee

Barroso, Antonio Carlos de Oliveira (President)
Baptista Filho, Benedito Dias
Bido, Diogenes de Souza

Title in Portuguese

Algoritmo rastreador web especialista nuclear

Keywords in Portuguese

algoritmos de busca e rastreamento web
informação nuclear na web
mineração de dados na web
recuperação de informações na web
sistemas especialistas nucleares

Abstract in Portuguese

Nos últimos anos a Web obteve um crescimento exponencial, se tornando o maior repositório de informações já criado pelo homem e representando uma fonte nova e relevante de informações potencialmente úteis para diversas áreas, inclusive a área nuclear. Entretanto, devido as suas características e, principalmente, devido ao seu grande volume de dados, emerge um problema desafiador relacionado à utilização das suas informações: a busca e recuperação informações relevantes e úteis. Este problema é tratado por algoritmos de busca e recuperação de informação que trabalham na Web, denominados rastreadores web. Neste trabalho é apresentada a pesquisa e desenvolvimento de um algoritmo rastreador que efetua buscas e recupera páginas na Web com conteúdo textual relacionado ao domínio nuclear e seus temas, de forma autônoma e massiva. Este algoritmo foi projetado sob o modelo de um sistema especialista, possuindo, desta forma, uma base de conhecimento que contem tópicos nucleares e palavras-chave que os definem e um mecanismo de inferência constituído por uma rede neural artificial perceptron multicamadas que efetua a estimação da relevância das páginas na Web para um determinado tópico nuclear, no decorrer do processo de busca, utilizando a base de conhecimento. Deste modo, o algoritmo é capaz de, autonomamente, buscar páginas na Web seguindo os hiperlinks que as interconectam e recuperar aquelas que são mais relevantes para o tópico nuclear selecionado, emulando a habilidade que um especialista nuclear tem de navegar na Web e verificar informações nucleares. Resultados experimentais preliminares apresentam uma precisão de recuperação de 80% para o tópico área nuclear em geral e 72% para o tópico de energia nuclear, indicando que o algoritmo proposto é efetivo e eficiente na busca e recuperação de informações relevantes para o domínio nuclear.

Title in English

Nuclear expert web crawler algorithm

Keywords in English

nuclear expert systems
nuclear web information
web mining
web retrieval
web search and crawling algorithms

Abstract in English

Over the last years the Web has obtained an exponential growth, becoming the largest information repository ever created and representing a new and valuable source of potentially useful information for several topics and also for nuclear-related themes. However, due to the Web characteristics and, mainly, because of its huge data volume, finding and retrieving relevant and useful information are non-trivial tasks. This challenge is addressed by web search and retrieval algorithms called web crawlers. This work presents the research and development of a crawler algorithm able to search and retrieve webpages with nuclear-related textual content, in autonomous and massive fashion. This algorithm was designed under the expert systems model, having, this way, a knowledge base that contains a list of nuclear topics and keywords that define them and an inference engine composed of a multi-layer perceptron artificial neural network that performs webpages relevance estimates to some knowledge base nuclear topic while searching the Web. Thus, the algorithm is able to autonomously search the Web by following the hyperlinks that interconnect the webpages and retrieving those that are more relevant to some predefined nuclear topic, emulating the ability a nuclear expert has to browse the Web and evaluate nuclear information. Preliminary experimental results show a retrieval precision of 80% for the nuclear general domain topic and 72% for the nuclear power topic, indicating that the proposed algorithm is effective and efficient to search the Web and to retrieve nuclear-related information.

WARNING - Viewing this document is conditioned on your acceptance of the following terms of use:
This document is only for private use for research and teaching activities. Reproduction for commercial use is forbidden. This rights cover the whole data about this document as well as its contents. Any uses or copies of this document in whole or in part must include the author's name.

2013ReisAlgoritmo.pdf (1.31 Mbytes)

Publishing Date

2014-01-14

Derived works

WARNING: Learn what derived works are clicking here.