Publications
A compact list of publications with links to the original research records and related resources.
Journal articles
Generalization Measures under Controlled Covariate Shift: A Regime-Aware Benchmark
TMLR 2026, August 2026 (accepted)
What do near-optimal learning rate schedules look like?
TMLR 2026, June 2026
Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training
TMLR 2025, August 2025
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
TMLR 2025, April 2025
Geometric Insights into Focal Loss: Reducing Curvature for Enhanced Model Calibration
PRL, February 2025
Towards Understanding Variants of Invariant Risk Minimization from the Perspective of Calibration
TMLR 2024, June 2024
Empirical Study on Optimizer Selection for Out-of-Distribution Generalizations
TMLR 2023, June 2023
Conference papers
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
ICML 2026, July 2026
On Fairness of Task Arithmetic: The Role of Task Vectors
ICLR 2026, January 2026
DiTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
ICLR 2026, January 2026
Mastering Task Arithmetic: τJp as a Key Indicator for Weight Disentanglement
ICLR 2025, April 2025
No Wrong Turns: The Simple Geometry Of Neural Networks Optimization Paths
ICML 2024, July 2024
How Image Corruption and Perturbation Affect Out-Of-Distribution Generalization and Calibration
IJCNN 2023, June 2023
Conjugate Gradient Method for Generative Adversarial Networks
AISTATS 2023, May 2023
Optimal Transport Meets Noisy Label Robust Loss and MixUp Regularization for Domain Adaptation
CoLLAs 2022, August 2022
Accelerating Convolutional Neural Networks Using Low Precision Arithmetic
HPC Asia 2018, January 2018
Accelerating Matrix Multiplication in Deep Learning by using Low-Rank Approximation
HPCS 2017, July 2017
Preprints and under review(7)
From Inner Randomness to Outer Stability
Under Review, September 2026 (under review)
Which Geometry on Which Layer? A Principled Criterion for Mixed-Optimizer Training
Under Review, May 2026 (under review)
Orth-Dion: Eliminating Geometric Mismatch in Distributed Low-Rank Spectral Optimization
Under Review, May 2026 (under review)
Convergence Bound and Critical Batch Size of Muon Optimizer
Under Review, July 2025 (under review)
When Does Alignment Help? A Comparative Study of DCCA and Fusion-Based Approaches for Multi-modal Chest X-ray Analysis
SSRN, April 2025
Augmenting NER Datasets with LLMs: Towards Automated and Refined Annotation
arXiv, March 2024
Takeuchi's Information Criteria as Generalization Measures for DNNs Close to NTK Regime
arXiv, September 2021
Workshop papers(9)
The Geometry of Spectral Gradient Descent: Layerwise Criteria for SignSGD vs SpecSGD
ICLR 2026 Workshop, February 2026
Smoothness-Adaptive Sharpness Aware Minimization for Finding Flatter Minima
ICLR 2024 Workshop, May 2024
Story-to-Images Translation: Leveraging Diffusion Models and Large Language Models for Sequence Image Generation
NarSUM 2023, October 2023
On the Interplay of Curvature, Calibration and Out-of-Distribution Generalization: Insights from SAM and Focal Loss Analyses
UNCV 2023, October 2023
Necessary and Sufficient Hypothesis of Curvature: Understanding Connection Between Out-of-Distribution Generalization and Calibration
ICLR 2023 Workshop, May 2023
Towards Understanding the Relationship of Batch Size and Iterations in Deep Learning
MLSS 2020, June 2020
On Empirical Analysis of Layer-wise Learning Rate Schedule
ACML 2019 Workshop, November 2019
A Performance Improvement Approach for Second-Order Optimization in Large Mini-batch Training
HPML 2019, May 2019
Noise Injection Leads to Better Generalization in Large Mini-Batch Training
Tokyo Tech-Stony Brook Joint Meeting, May 2019
Equal contribution