Using data mining methods to remove uncertainties in sensor data streams. This project will develop key techniques for removing uncertainties in sensor data streams and thus improve the monitoring quality of sensor networks. The expected outcomes will benefit Australia by enabling improved, lower-cost monitoring of natural resources and management of stock raising.
Deep Data Mining for Anomaly Prediction from Sensor Data Streams. Sensor data streams are crucial for anomaly predictions in real-life monitoring. However, balancing efficiency and accuracy in predicting anomalies with sensor streams is a great challenge; it requires new techniques that go beyond detecting anomalies and predicting trends. This project will develop a deep mining method for anomaly prediction from sensor streams; it will comprise mining algorithms at various levels - from compress ....Deep Data Mining for Anomaly Prediction from Sensor Data Streams. Sensor data streams are crucial for anomaly predictions in real-life monitoring. However, balancing efficiency and accuracy in predicting anomalies with sensor streams is a great challenge; it requires new techniques that go beyond detecting anomalies and predicting trends. This project will develop a deep mining method for anomaly prediction from sensor streams; it will comprise mining algorithms at various levels - from compressing massive raw data, to recognition of abnormal waveforms preceding anomalies, and to retrieving and summarising similar past anomalies for creating descriptions of future anomalies. The project will demonstrate our method in health/environment monitoring applications, and its adoption will save resources, money and lives.Read moreRead less
Knowledge discovery from data in the context of prior beliefs. This project will invent user-centric technologies for discovering knowledge from data that are distinguished by taking account of the user's beliefs, enabling more useful discoveries to be made. This project will also invent methods that identify the relative potential value of those discoveries, helping the user derive greater value from their data assets.
Combining generative and discriminative strategies to facilitate efficient and effective learning from big data. Effective extraction of information from massive data stores is increasingly problematic as data quantities continue to grow rapidly. Quite simply, effective techniques for learning from small data do not scale. However, the problem is even worse than this. Big data contain more information than the small data in which context most state-of-the-art learning algorithms have been develo ....Combining generative and discriminative strategies to facilitate efficient and effective learning from big data. Effective extraction of information from massive data stores is increasingly problematic as data quantities continue to grow rapidly. Quite simply, effective techniques for learning from small data do not scale. However, the problem is even worse than this. Big data contain more information than the small data in which context most state-of-the-art learning algorithms have been developed. For small data overly detailed classifiers will overfit the data and so should be avoided. In contrast, big data provide fine detail and hence will benefit new types of learner that can capture it. This project will deliver learners that are not only capable of capturing this detail, but do so with the efficiency required to process terabytes of data.Read moreRead less
Mining large negative correlations for high-dimensional contrasting analysis. Negative correlations are widely embedded in real life applications, but in-depth research has rarely been conducted due to its high level of complexity. This project aims at efficient algorithms and frontier theory for finding large negative correlations, to enable smart information use in bioinformatics to promote Australia's leading role in data mining research.
Efficient causal discovery from observational data. Discovering cause-effect relationships is the ultimate goal for many applications. Randomised control trial is the gold standard for discovering causal relationships. However, conducting such trials is impossible in many cases due to cost and/or ethical concerns. In contrast, a large amount of data has been accumulated in all areas. It is desirable to infer causal relationships from data directly and automatically. This project aims to develop ....Efficient causal discovery from observational data. Discovering cause-effect relationships is the ultimate goal for many applications. Randomised control trial is the gold standard for discovering causal relationships. However, conducting such trials is impossible in many cases due to cost and/or ethical concerns. In contrast, a large amount of data has been accumulated in all areas. It is desirable to infer causal relationships from data directly and automatically. This project aims to develop fast and scalable data mining methods for identifying causal relationships from large and/or high dimensional data sets. The developed methods will mainly be evaluated in real world biological applications. The research outcomes will be useful in many areas for causal reasoning and decision making.Read moreRead less
Developing novel data mining methods to reveal complex group relationships from heterogeneous data. This project aims to develop novel and effective data mining methods that will enable us to unravel the relationships between multiple, rather than individual, components of complex systems (such as genes, gene regulators and cancer), which is crucial to understanding how such systems work. Potential applications for such methods are extensive.
Online Learning for Large Scale Structured Data in Complex Situations. Online Learning (OL) is the process of predicting answers for a sequence of questions. OL has enjoyed much attention in recent years due to its natural ability of processing large scale non-structured data and adapting to a changing environment. However, OL has three weaknesses: it does not scale for structured data; it often assumes that all of the data are equally important; it often considers that all of the data are compl ....Online Learning for Large Scale Structured Data in Complex Situations. Online Learning (OL) is the process of predicting answers for a sequence of questions. OL has enjoyed much attention in recent years due to its natural ability of processing large scale non-structured data and adapting to a changing environment. However, OL has three weaknesses: it does not scale for structured data; it often assumes that all of the data are equally important; it often considers that all of the data are complete and noise-free. These weaknesses limit its utility, because real data such as those that must be analysed in processing social networks, fraud detection do not satisfy the restrictions. The aim of this project is to develop theoretical and practical advances in OL that overcome the existing weaknesses.Read moreRead less
Probabilistic Graphical Models For Interventional Queries. The project intends to develop methods to suggest how to optimally intervene so that the future state of the system will best suit our interests. The power of probabilistic graphical models to model complex relationships and interactions among a large number of variables facilitates many applications. However, such models only aim to understand the underlying environment. What is ultimately needed in many real-world applications is to su ....Probabilistic Graphical Models For Interventional Queries. The project intends to develop methods to suggest how to optimally intervene so that the future state of the system will best suit our interests. The power of probabilistic graphical models to model complex relationships and interactions among a large number of variables facilitates many applications. However, such models only aim to understand the underlying environment. What is ultimately needed in many real-world applications is to suggest how we ought to intervene or act, so as to alter the environment to best suit our interests. The proposed project aims to achieve this using probabilistic graphical models on massive real-world data sets, thus facilitating a variety of applications from health care to commerce and the environment.Read moreRead less
Coupling Learning in Big Data. Big data features complex coupling relationships within and between diverse entities in various forms and layers. This fundamentally challenges existing learning theories, which usually assume that data is independent and identically distributed (IID). This indicates that such IID tools may either be inapplicable for big data or capture an incomplete or even biased picture of the ground truth in big data. Hence, this project aims to invent breakthrough theories and ....Coupling Learning in Big Data. Big data features complex coupling relationships within and between diverse entities in various forms and layers. This fundamentally challenges existing learning theories, which usually assume that data is independent and identically distributed (IID). This indicates that such IID tools may either be inapplicable for big data or capture an incomplete or even biased picture of the ground truth in big data. Hence, this project aims to invent breakthrough theories and effective tools for systematically modelling and learning sophisticated couplings embedded in big data applications. The outcomes are expected to enhance Australia's leading role in data science research and lift data intelligence-driven productivity and economic growth in a changing world.Read moreRead less