Indexes Allowing Fast and Efficient Text Search. Since the arrival of search engines such as Google, it has become an expectation that we can find a few words in a very large amount of text very quickly. This is also true of biologists, who expect to be able to submit protein sequences for matching against a massive database and retrieve and answer in seconds. If successful, this project will invent fundamental software that will allow the discovery of words in text much more quickly, and using ....Indexes Allowing Fast and Efficient Text Search. Since the arrival of search engines such as Google, it has become an expectation that we can find a few words in a very large amount of text very quickly. This is also true of biologists, who expect to be able to submit protein sequences for matching against a massive database and retrieve and answer in seconds. If successful, this project will invent fundamental software that will allow the discovery of words in text much more quickly, and using less computing resources than current methods. The benefits of this technology to the searching public, scientists, and industry will be immediate as productivity will be improved and costs reduced.Read moreRead less
New approaches to interactive sessional search for complex tasks. This project aims to develop new tools and techniques to improve the accuracy and speed of search and data analytics for complex information tasks. There are currently no publicly available search engines which support users engaged in complex interactive search, or that allow searchers to fully control their own data and privacy. Fundamental research advances, based on understanding real user behaviour and search needs will have ....New approaches to interactive sessional search for complex tasks. This project aims to develop new tools and techniques to improve the accuracy and speed of search and data analytics for complex information tasks. There are currently no publicly available search engines which support users engaged in complex interactive search, or that allow searchers to fully control their own data and privacy. Fundamental research advances, based on understanding real user behaviour and search needs will have an impact on important academic, industrial, and government domains, including virtual assistants, health care (clinical decision support), precision medicine, eDiscovery, crime prevention, and detailed socio-economic evaluations.Read moreRead less
Automatic music feature extraction, classification and annotation. Music is a huge industry currently undergoing a major revolution. The industry is shifting from music-making to music retrieval and its incorporation into a range of products from TV, and film, to music streaming into locations and events, as well as MP3 players and all kinds of electronic devices. This research will support immediate retrieval of music that meets the current industry need, based not just on titles, composers and ....Automatic music feature extraction, classification and annotation. Music is a huge industry currently undergoing a major revolution. The industry is shifting from music-making to music retrieval and its incorporation into a range of products from TV, and film, to music streaming into locations and events, as well as MP3 players and all kinds of electronic devices. This research will support immediate retrieval of music that meets the current industry need, based not just on titles, composers and/or performers, but on the actual properties of the music itself. The knowledge and music processing techniques developed will give Australian music industry an advantage over other countries.Read moreRead less
Development and Application of Techniques for Detecting Equivalent Documents. The web is a vast collection of data, such as text and images, but contains large numbers of duplicates - the same document or picture may be present many times. Even personal collections of information, such as the documents and digital photos people keep on their home computers, often have many versions of the same item. However, detecting such duplicates is not straightforward, as they may have been edited, or may, ....Development and Application of Techniques for Detecting Equivalent Documents. The web is a vast collection of data, such as text and images, but contains large numbers of duplicates - the same document or picture may be present many times. Even personal collections of information, such as the documents and digital photos people keep on their home computers, often have many versions of the same item. However, detecting such duplicates is not straightforward, as they may have been edited, or may, for example, be shown in different forms; for example, the quality of a photo may be reduced for display on a mobile phone. In this project we plan to detect such duplicates, and use the results to improve search and management of data.Read moreRead less
Effective Information Retrieval for Partitioned Document Collections. Current information retrieval services make use of massive indexes in order to resolve content-based queries. Monolithic approaches like this have been effective until now because the volume of data stored has been manageable on a single machine or tightly-coupled cluster of machines, and because the data has been available for collection. But with an increasing amount of automatically generated data, and an increasing diversi ....Effective Information Retrieval for Partitioned Document Collections. Current information retrieval services make use of massive indexes in order to resolve content-based queries. Monolithic approaches like this have been effective until now because the volume of data stored has been manageable on a single machine or tightly-coupled cluster of machines, and because the data has been available for collection. But with an increasing amount of automatically generated data, and an increasing diversity of information sources, other approaches are required. In this project we will investigate mechanisms for handling retrieval tasks when the indexes to the data are stored locally with the data, and when no central index is viable.Read moreRead less
Dynamic Index Maintenance for Text Search Engines. Text retrieval systems such as internet search engines use high-performance indexes to rapidly locate documents that match user queries. In recent years there have been major improvements in query evaluation and index construction techniques. As the data changes, it is necessary to keep the index up to date, but current methods for maintaining indexes are slow and costly. The aim of this project is to develop methods that provide on-the-fly u ....Dynamic Index Maintenance for Text Search Engines. Text retrieval systems such as internet search engines use high-performance indexes to rapidly locate documents that match user queries. In recent years there have been major improvements in query evaluation and index construction techniques. As the data changes, it is necessary to keep the index up to date, but current methods for maintaining indexes are slow and costly. The aim of this project is to develop methods that provide on-the-fly update at much lower cost, thereby improving the performance of text retrieval systems. This work involves both practical development and innovation in fundamental algorithms.Read moreRead less