lensen2021mining.pdf (542.78 kB)

Mining Feature Relationships in Data

Download (542.78 kB)
conference contribution
posted on 12.08.2021, 03:25 by Andrew LensenAndrew Lensen
When faced with a new dataset, most practitioners begin by performing exploratory data analysis to discover interesting patterns and characteristics within data. Techniques such as association rule mining are commonly applied to uncover relationships between features (attributes) of the data. However, association rules are primarily designed for use on binary or categorical data, due to their use of rule-based machine learning. A large proportion of real-world data is continuous in nature, and discretisation of such data leads to inaccurate and less informative association rules. In this paper, we propose an alternative approach called feature relationship mining (FRM), which uses a genetic programming approach to automatically discover symbolic relationships between continuous or categorical features in data. To the best of our knowledge, our proposed approach is the first such symbolic approach with the goal of explicitly discovering relationships between features. Empirical testing on a variety of real-world datasets shows the proposed method is able to find high-quality, simple feature relationships which can be easily interpreted and which provide clear and non-trivial insight into data.

History

Preferred citation

Lensen, A. (2021, January). Mining Feature Relationships in Data. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (12691 LNCS pp. 247-262). Springer International Publishing. https://doi.org/10.1007/978-3-030-72812-0_16

Title of proceedings

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)

Volume

12691 LNCS

Publication or Presentation Year

01/01/2021

Pagination

247-262

Publisher

Springer International Publishing

Publication status

Published

ISSN

0302-9743

eISSN

1611-3349

Usage metrics

Read the peer-reviewed publication

Categories

Exports