Skip to main content
WHITE PAPER

The importance of data selection in AI model training, testing, and deployment

5 October 2026

The successful training and deployment of artificial intelligence (AI) tools are highly dependent on the quality of their input data. As the proliferation of AI tools for healthcare applications continues, it is critically important that stakeholders consider the provenance and quality of any data used to power an AI application. Healthcare AI model output success depends on data selection, for small differences in initial inputs may produce vastly different outcomes. There are many different types and flavors of healthcare data, each suitable for some tasks and not for others, making data selection difficult. In this overview, we describe some of the considerations at play in bringing the right data to bear for a given use.

Our discussion points include the following.

  • Data confidence: Selecting the right data
  • Variance in healthcare claims data: Examining open, closed, and paid claims
  • Reference data: Providing additional context by reviewing entities, geographies, or concepts involved in claims and enrollment data

Download the overview (PDF).


Explore more tags from this article

About the Author(s)

Jeff McGinn

Robert Richards

We’re here to help