About me
I'm Shuyi Fan — also known as Fancy — and I work on how AI tutoring systems are evaluated, and whether the measures the field relies on actually capture teaching. My first-author paper audits a common assumption: that a general-purpose "helpfulness" rating can distinguish a tutor that guides a student from one that simply hands over the answer. Under a controlled, pre-registered design, it can't — the ordering flips depending on which model is doing the judging.
I hold an M.A. in Communication and Education from Columbia University and a B.S. in Media, Culture, and Communication from New York University. That training is why I read evaluation as more than a benchmarking problem: deciding what counts as good teaching is a measurement question and a learning question at once, and the two don't come apart cleanly.
Alongside research I co-founded Dr. Milou (Paw Paw Inc.), an at-home pet care venture built on AI tools and smart hardware, and earlier worked as an analyst at Nielsen, Uber (Hong Kong), and Lions Financial — designing and interpreting qualitative and quantitative studies, including a survey of more than 1,000 respondents. I'm looking for doctoral work on AI and education that takes measurement seriously.
Research interests
-
LLM Tutoring & Pedagogy
What separates a tutor that teaches from one that answers, and whether current systems can tell the difference.
-
Evaluation & LLM-as-Judge
Where model-based evaluation is load-bearing, where the choice of judge quietly decides the result, and what a rubric can't see.
-
Measuring Learning Processes
Behavioral traces — answer leakage, independent work on the next turn — as evidence that judged ratings alone can't supply.
-
Reproducible Study Design
Pre-registration, condition-blind judging, and robustness audits that keep a finding standing when the judge changes.