Research Interests
I work on how we tell whether a language model is actually good at something, and on what happens when a person and a model have to reach an answer together rather than one replacing the other — evaluation reliability, and human–AI complementarity. Before that, computer vision: video understanding and annotation, and interfaces that give people real-time feedback on their own performance.
Selected Publications
- Toward Human-AI Complementarity Across Diverse Tasks Under review · arXiv:2605.04070
- Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge Under review · arXiv:2602.02219
- MK2 at PBIG Competition: A Prompt Generation Solution IJCAI AgentScen Workshop 2025 · ACL Anthology · arXiv
- VisDev: A Stroke Deviation Visualization System Using Point-Set Registration for Handwriting Training CHI Late-Breaking Work 2025 · ACM DL
- Video Region Annotation with Sparse Bounding Boxes IJCV 2023 · BMVC 2020 (Best Student Paper) · Springer · arXiv