I am Jirui Dai (戴纪瑞), a Master’s student in Computer Science at Johns Hopkins University.
I am currently working across multiple research groups:
Assistant Researcher (monthly stipend) in the Cao Peng Group at Nanjing University of Chinese Medicine, advised by Postdoctoral Researcher Zhi Liu, with collaborative research involving multiple medical institutions (including Peking Union Medical College Hospital, Dongfang Hospital of Beijing University of Chinese Medicine, and Gulou Hospital of Traditional Chinese Medicine of Beijing et al).
Advised by Prof. Mark Dredze, collaborating with Ph.D. student Heyuan Huang.
My interests sit at the intersection of Medical AI and Reinforcement Learning, with a focus on building clinically grounded and capability-expanding AI systems.
- LLM / MLLM / Agent for clinical practice: turning expert knowledge and real-world medical signals into reliable, usable systems—emphasizing structured reasoning, robustness, and evaluation that reflects clinical reality.
- RL for foundation models: exploring how reinforcement learning can expand model capabilities beyond “alignment,” toward broader and more reliable solution spaces.
I’m actively looking to collaborate with researchers working on medical AI, multimodal learning, or RL for foundation models.
You can find my CV here (updated through 01/15/2026).
Research Taste
I am cautiously optimistic about AGI: progress is real, but confident narratives rarely survive stress testing.
My strongest conviction lies in healthcare AI. Medicine faces structural scarcity — limited expert time, uneven access, high cognitive load — and carefully designed AI can help scale expertise and broaden access, provided we treat evaluation, safety, and distribution shift as first-class concerns.
On the methods side, I am drawn to reinforcement learning. RL captures something fundamental: real-world decisions involve long chains of trade-offs with implicit rewards. I believe there is latent structure in the incentive landscape that can steer foundation models toward a wider output manifold — richer strategies, more robust reasoning, and more creative problem-solving — rather than collapsing into narrow, overly safe modes.
