Efficient and Effective Text Information Retrieval with Phrases. Current Internet search engines find documents by matching queries to documents, then present the closest matches to the user. Such searching is often ineffective. Another technique for searching, which with current algorithms is not feasible for large text collections such as the Web, is to browse vocabularies and view phrases in the contexts in which they are used. The aim of this project is to make wider use of phrases in re ....Efficient and Effective Text Information Retrieval with Phrases. Current Internet search engines find documents by matching queries to documents, then present the closest matches to the user. Such searching is often ineffective. Another technique for searching, which with current algorithms is not feasible for large text collections such as the Web, is to browse vocabularies and view phrases in the contexts in which they are used. The aim of this project is to make wider use of phrases in retrieval, by developing new phrase-based querying algorithms and investigating how users can use phrase indexes to find documents. The outcome will be new, efficient methods for exploring the Internet.Read moreRead less
Discovery Early Career Researcher Award - Grant ID: DE140100275
Funder
Australian Research Council
Funding Amount
$392,979.00
Summary
Beyond keyword search for ranked document retrieval. This project will develop novel approaches to efficient and effective ranked text retrieval using a new class of rank-aware algorithms derived from self-indexes. These algorithms can support complex statistical calculations on the fly. Efficient algorithm design for big data is an increasingly important problem as energy costs continue to soar and can now exceed hardware costs for big data consumers such as Google. In this project, two importa ....Beyond keyword search for ranked document retrieval. This project will develop novel approaches to efficient and effective ranked text retrieval using a new class of rank-aware algorithms derived from self-indexes. These algorithms can support complex statistical calculations on the fly. Efficient algorithm design for big data is an increasingly important problem as energy costs continue to soar and can now exceed hardware costs for big data consumers such as Google. In this project, two important problems in web search are explored: real-time indexing and long-form query answering. Using self-index algorithms, this project presents a road map to move beyond simple keyword-based ranked document retrieval, thus allowing us to efficiently meet more demanding information needs of users in the next decade.Read moreRead less
Fast and Scalable Search Techniques for Genomic Databases. Tens of thousands of users each day search the genomic databases that are so far the most significant product of the Human Genome Project. In this project, we will investigate fundamental new bioinformatics techniques for information retrieval from genomic databases. The outcomes will allow molecular biologists to accurately and efficiently discover relationships between DNA and protein sequences. In contrast to existing approaches, o ....Fast and Scalable Search Techniques for Genomic Databases. Tens of thousands of users each day search the genomic databases that are so far the most significant product of the Human Genome Project. In this project, we will investigate fundamental new bioinformatics techniques for information retrieval from genomic databases. The outcomes will allow molecular biologists to accurately and efficiently discover relationships between DNA and protein sequences. In contrast to existing approaches, our techniques will remain fast despite the enormous growth in genomic database sizes. This research will contribute significantly to the "key areas of study includ[ing] genomics and bioinformatics" in the new ARC genome/phenome link priority area.Read moreRead less
Using Past Queries for Fast and Accurate Web Searching. Searching the entire Internet, or a company web site, has become a vital task for modern organisations. While there has been significant research into improving search engines through using web pages themselves, very little attention has been paid to improving web search by exploiting the vast numbers of queries that users submit to search engines each day. This project will use state of the art compression and algorithmic techniques to imp ....Using Past Queries for Fast and Accurate Web Searching. Searching the entire Internet, or a company web site, has become a vital task for modern organisations. While there has been significant research into improving search engines through using web pages themselves, very little attention has been paid to improving web search by exploiting the vast numbers of queries that users submit to search engines each day. This project will use state of the art compression and algorithmic techniques to improve the speed and accuracy of web search using data gleaned from millions of Internet queries (provided under agreement by Microsoft). Improving search engines will have a direct benefit to many Australian industries, and support the government's priority area of "smart information use".Read moreRead less
Spoken conversational search: contextual interactive techniques to support effective information search over a speech-only communication channel. This project will develop new techniques for effective information search using speech only, supporting improved information access for visually impaired people or in situations that require focused visual attention (e.g. driving). The techniques are based on a conversational approach to information search and presentation of results.
Development and Application of Techniques for Detecting Equivalent Documents. The web is a vast collection of data, such as text and images, but contains large numbers of duplicates - the same document or picture may be present many times. Even personal collections of information, such as the documents and digital photos people keep on their home computers, often have many versions of the same item. However, detecting such duplicates is not straightforward, as they may have been edited, or may, ....Development and Application of Techniques for Detecting Equivalent Documents. The web is a vast collection of data, such as text and images, but contains large numbers of duplicates - the same document or picture may be present many times. Even personal collections of information, such as the documents and digital photos people keep on their home computers, often have many versions of the same item. However, detecting such duplicates is not straightforward, as they may have been edited, or may, for example, be shown in different forms; for example, the quality of a photo may be reduced for display on a mobile phone. In this project we plan to detect such duplicates, and use the results to improve search and management of data.Read moreRead less
Effective Information Retrieval for Partitioned Document Collections. Current information retrieval services make use of massive indexes in order to resolve content-based queries. Monolithic approaches like this have been effective until now because the volume of data stored has been manageable on a single machine or tightly-coupled cluster of machines, and because the data has been available for collection. But with an increasing amount of automatically generated data, and an increasing diversi ....Effective Information Retrieval for Partitioned Document Collections. Current information retrieval services make use of massive indexes in order to resolve content-based queries. Monolithic approaches like this have been effective until now because the volume of data stored has been manageable on a single machine or tightly-coupled cluster of machines, and because the data has been available for collection. But with an increasing amount of automatically generated data, and an increasing diversity of information sources, other approaches are required. In this project we will investigate mechanisms for handling retrieval tasks when the indexes to the data are stored locally with the data, and when no central index is viable.Read moreRead less
Dynamic Index Maintenance for Text Search Engines. Text retrieval systems such as internet search engines use high-performance indexes to rapidly locate documents that match user queries. In recent years there have been major improvements in query evaluation and index construction techniques. As the data changes, it is necessary to keep the index up to date, but current methods for maintaining indexes are slow and costly. The aim of this project is to develop methods that provide on-the-fly u ....Dynamic Index Maintenance for Text Search Engines. Text retrieval systems such as internet search engines use high-performance indexes to rapidly locate documents that match user queries. In recent years there have been major improvements in query evaluation and index construction techniques. As the data changes, it is necessary to keep the index up to date, but current methods for maintaining indexes are slow and costly. The aim of this project is to develop methods that provide on-the-fly update at much lower cost, thereby improving the performance of text retrieval systems. This work involves both practical development and innovation in fundamental algorithms.Read moreRead less
Smart Algorithms Linking Medical Image Data and Measures of Dysfunction. Losing sight has a profound affect on a person's quality of life. Advances in devices that monitor vision have not been matched by advances in computer software that analyses data from those devices. This project will combine computer science, visual neuroscience and clinical expertise to devise algorithms and build software that will vastly improve clinician's abilities to diagnose and monitor vision loss. In turn, this wi ....Smart Algorithms Linking Medical Image Data and Measures of Dysfunction. Losing sight has a profound affect on a person's quality of life. Advances in devices that monitor vision have not been matched by advances in computer software that analyses data from those devices. This project will combine computer science, visual neuroscience and clinical expertise to devise algorithms and build software that will vastly improve clinician's abilities to diagnose and monitor vision loss. In turn, this will dramatically improve the chances of those with diseases such as glaucoma to preserve their sight into old age. Furthermore, outcomes from this project will inform the development bionic eye technologies, which will assist those with eye diseases such as retinis pigmantosa and age-related macular degeneration to see.Read moreRead less