ICML2024
Day1 (7/22 Mon)
-
Towards Efficient Generative Large Language Model Serving: A Tutorial from Algorithms to Systems
- arxiv
- speaker: Xupeng Miao
- from the system research area
- overview
- Various methods are updated daily, mainly to improve LLM inference (computational cost, token scale, and quality).
- Systematic explanation of the recent evolution of algorithms and systemic aspects
- thought
- It is important to be able to organize taxonomy in any field (otherwise, it is impossible to cite correct papers and to create correct research directions).
- It is also important for both practical and research purposes to know how to make good use of current methods (MoE, LoRA, Continual Learning, etc.) under the constraints of numerous speed-up and efficiency-enhancing methods.
- e.g. QLoRA
- details related to my interest
-
Lunch w/ researchers from RIKEN
Day2 (7/23 Tue)
-
Opening Remarks
- Acceptance rate: 27%
- trend: LLM
-
Coffee break
-
Poster Session
- meet
- work
- WIP
- tons of interesting works
- WIP
-
Lunch
Day3 (7/24 Wed)
-
Our poster presentation (13:00-)
- No Wrong Turns: The Simple Geometry Of Neural Networks Optimization Paths
- ICML portal
- paper
- I felt sick and stayed in my room almost all day.
- I was diagnosed as COVID-19 positive.
Day4 (7/25 Thu)
- Stay Hotel due to COVID-19
Day5 (7/26 Fri)
- Stay Hotel due to COVID-19
Papers
Industry
Google DeepMind
-
papers
- Information Complexity of Stochastic Convex Optimization: Applications to Generalization and Memorization
- LEVI: Generalizable Fine-tuning via Layer-wise Ensemble of Different Views
- Interpretability Illusions in the Generalization of Simplified Models
- Position: Leverage Foundational Models for Black-Box Optimization
- With Sakana.AI
Microsoft Research
-
papers
- Precise Accuracy / Robustness Tradeoffs in Regression: Case of General Norms
- Tag-LLM: Repurposing General-Purpose LLMs for Specialized Domains
- Selective Mixup Helps with Distribution Shifts, But Not (Only) because of Mixup
- Author: Damien Teney
- DoRA: Weight-Decomposed Low-Rank Adaptation
- Open-Vocabulary Calibration for Vision-Language Models
- Breaking through the learning plateaus of in-context learning in Transformer
Apple MLR
-
papers
- Optimization Without Retraction on the Random Generalized Stiefel Manifold
- Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
- Careful With That Scalpel: Improving Gradient Surgery With an EMA
- Revealing the Utilized Rank of Subspaces of Learning in Neural Networks
Amazon Science
-
papers
- MADA: Meta-adaptive optimizers through hyper-gradient descent
- Multicalibration for confidence scoring in LLMs
- Transferring knowledge from large foundation models to small downstream models
- Using uncertainty quantification to characterize and improve out-of-domain learning for PDEs
- Variance-reduced zeroth-order methods for fine-tuning language models
Acknowlegements
I want to thank my PhD supervisor Ioannis Mitliagkas (Mila, UdeM) for supporting my participation in the ICML.