Michael Ginn
/maɪkəl dʒɪn/
I'm a fifth-year Ph.D. candidate at the University of Colorado studying natural language processing, machine learning, and occasionally linguistics, supervised by Prof. Alexis Palmer and Prof. Mans Hulden.
My Research
I am broadly interested in multilinguality in language models, low-resource settings, synthetic data, reinforcement learning, and self-distillation. I have done a lot of research involving training multilingual models for endangered language documentation. These include GlossLM in 2024 and PolyGloss in 2026, the latter winning an Outstanding Paper award at ACL.
I have interned a number of times at Apple in ML engineering roles, most recently training a next-edit prediction model for coding entirely on synthetic data. Right now, I'm an Applied Science Intern at Amazon AWS AI, where I'm studying crosslingual self-distillation in vision-language models.
Some big questions I'm thinking about right now:
- How can we make models as multilingual as possible (i.e. equally capable in a huge number of languages) while retaining language-specific information and maximal capabilities?
- How are languages, particularly low-resource languages, represented across domains and tasks in a language model? Are there specific characteristics we can use to predict and correct crosslingual discrepencies?
- Is true continual learning (i.e. both knowledge updates and increasing intelligence) possible with the current language modeling paradigm? One possibility is self-distillation methods, but are these methods really capable of meaningful and stable improvements?
- How and when is synthetic data (particularly that created without use of a stronger teacher) effective? Can we simulate data in a way that adds new information to learn from, or is there a fundamental limitation?