Adaptive Generation of Training Data for ML Reduced Model Creation

Show authors

Publication Type

Conference Paper

Book Title

2022 IEEE International Conference on Big Data (Big Data)

Publication Date

December, 2022

Page Numbers

3408 to 3416

Publisher Location

New Jersey, United States of America

Conference Name

The 4th International Workshop on Big Data Tools, Methods, and Use Cases for Innovative Scientific Discovery (BTSD) 2022

Conference Location

Virtual, Tennessee, United States of America

Conference Sponsor

IEEE

Conference Date

Dec 17, 2022 - Dec 20, 2022

View DOI Listing

Abstract

Machine learning proxy models are often used to speed up or completely replace complex computational models. The greatly reduced and deterministic computational costs enable new use cases such as digital twin control systems and global optimization. The challenge of building these proxy models is generating the training data. A naive uniform sampling of the input space can result in a non-uniform sampling of the output space of a model. This can cause gaps in the training data coverage that can miss finer scale details resulting in poor accuracy. While larger and larger data sets could eventually fill in these gaps, the computational burden of full-scale simulation codes can make this prohibitive. In this paper, we present an adaptive data generation method that utilizes uncertainty estimation to identify regions where training data should be augmented. By targeting data generation to areas of need, representative data sets can be generated efficiently. The effectiveness of this method will be demonstrated on a simple one-dimensional function and a complex multidimensional physics model.

Adaptive Generation of Training Data for ML Reduced Model Creation

Abstract

Researchers

Organizations