
Ponente: Diego Furtado Silva (University of Sao Paulo)
Resumen / Abstract:
"In the standard machine learning pipeline, we usually focus on assigning individual labels to new data and optimizing our models for instance-level accuracy. However, in many real-world scenarios, from monitoring public health trends to analyzing large-scale consumer sentiment, we are interested in summarizing the overall prevalence of classes within an unlabeled batch of data. This talk introduces the paradigm of "learning to quantify" (a.k.a. prevalence estimation or quantification), an essential but often overlooked field dedicated to accurately estimating class distributions within populations. We will explore why even the most sophisticated classifiers often fail as estimators when faced with label shift and how this discrepancy can lead to costly miscalculations in business and research. Moving beyond simple binary scenarios, we will discuss how to leverage this paradigm for structured data, specifically focusing on hierarchical relationships where labels are nested. By shifting the focus from "Who is this instance?" to "What is the true prevalence of this group?", attendees will gain a new perspective on model evaluation and learn practical strategies to derive reliable, high-level insights from noisy, real-world predictions, emphasizing the importance of population-level understanding for impactful research and decision-making."