Research

통계 방법론과 과학·사회문제의 데이터 분석을 연구합니다.

Research Overview

We develop statistical methodologies for understanding complex scientific and social phenomena and addressing real-world problems. We are particularly interested in developing new methods of statistical inference and computational tools that can effectively handle the dependence, uncertainty, and complex structure encountered in high-dimensional, large-scale data.

We collaborate with researchers across a wide range of fields—including astronomy, particle physics, neuroscience, biomechanics, epidemiology, digital humanities, and computational social science—to reformulate scientific questions as statistical problems and answer them through reproducible analysis. Our main research areas are described below.

Multiple Testing and False Discovery Rates

We study problems that require testing a large number of hypotheses simultaneously, as arise in genomics, imaging, and large-scale experimental data. Our recent work focuses on multiple-testing procedures and Cauchy combination tests that remain reliable even when test statistics exhibit complex or unknown dependence. We also develop semi-parametric methods for estimating local false discovery rates (local FDRs) by combining information across multiple studies and datasets. In addition to theoretical methodology, we pursue software and reproducible analytical tools that can be used by applied researchers.

Modern Dependence Measures

Relationships among variables often cannot be adequately described by linear correlation alone. High-dimensional and complex scientific data may exhibit nonlinear, local, and asymmetric dependence structures. We study modern dependence measures that can effectively detect and quantify these diverse relationships. We investigate their theoretical properties and associated inferential procedures, with the goal of applying them to independence testing, variable selection, clustering, and the exploration of complex data structures.

Sequential Testing and Quality Control

In steel and semiconductor manufacturing, high-frequency sensor readings, functional profiles, images, and inspection results accumulate continuously, while quality labels may arrive late or be observed only partially. Moreover, data distributions can change over time because of changes in products, equipment, maintenance, or raw materials. In these settings, it is important not only to achieve predictive accuracy at a single point in time, but also to ensure that false alarms do not accumulate during repeated monitoring, predictive uncertainty is properly assessed, and root-cause analysis following an alert is reliable.

We study reliable inference for data streams in which dependence, distributional shifts, missingness, and delayed labels coexist. Our work combines sequential testing, e-values and e-processes, change-point detection, online conformal prediction, structured reduction, and post-selection inference. Our goal is to build a layer of statistical reliability on top of conventional statistical process-control or AI-based anomaly-detection models—controlling alert errors regardless of when monitoring stops and consistently quantifying uncertainty throughout prediction, alerting, and root-cause investigation.

Our primary applications are quality control in steel and semiconductor manufacturing. We jointly evaluate prediction intervals, anomaly alerts, and error rates for candidate root causes using steel-process thickness and shape profiles, surface defects, and semiconductor equipment sensors, virtual metrology, wafer maps, and inspection images. We aim to implement our research through reproducible simulation studies, public benchmarks, and R/Python software.

Shrinkage Estimation and Statistical Learning

In high-dimensional settings, individual models and estimators can become unstable under small changes in the data. To address this issue, we study James–Stein shrinkage estimation, model aggregation, bagging, and ensemble methods.

Our central objective is to effectively combine information from multiple candidate estimators and predictive models, improving predictive accuracy and stability while enabling reliable statistical decision-making in settings with complex data structures.

Functional Data Analysis and Biomechanics

Movements, postures, forces, and joint trajectories measured continuously over time should often be understood not as individual observations, but as functions or curves. Our laboratory uses functional data analysis to study the complex patterns in such biomechanical data.

We develop methods to compare individual and group differences during movement, extract dominant movement patterns, and identify functional features associated with disease, injury, and training effects. Through functional classification, dimension reduction, registration, clustering, and uncertainty quantification, we aim to provide interpretable and reproducible statistical evidence for biomechanics research.

Spatial Statistics

We study data for which spatial location and interaction are essential, including the distribution of stars and galaxies, ecological and environmental data, and regional population data. Our work develops statistical models that estimate intensity functions and correlation structures for spatial point processes, select bandwidths, and account for regional heterogeneity and dependence.

Statistical Applications to Scientific and Social Problems

We seek to apply new statistical methods to real-world problems and to turn methodological questions that arise in applications into new research directions. Our recent work addresses diverse topics, including particle physics; analysis of constitutional court justices’ ideological tendencies and wildlife population estimation.

Opportunities for Students

Our laboratory is seeking undergraduate research interns and master’s and doctoral students who are interested in statistical methodology and data-driven problem solving.

Undergraduate research interns can gain experience with the full research process under supervision, including reading papers, data analysis, simulation, and programming. Graduate students can develop research topics aligned with their interests in multiple testing, statistical inference for particle physics, modern dependence measures, sequential testing, conformal prediction, change-point detection, quality control, statistical learning, functional data analysis, biomechanics, or spatial statistics.

We welcome students who wish to study new methods in depth and take on the compelling questions posed by real data. Students interested in research participation or graduate study are encouraged to contact us by email with a brief self-introduction and their research interests.