Hi! I’m Artemis.
I’m an M.S. student in Computer Science (Artificial Intelligence) at Stanford University, and an engineer in Apple’s Silicon Engineering Group. Previously, I earned my B.S.E. in Electrical and Computer Engineering (Computer Systems) from Princeton University. My work and research focus on quantitative methods for understanding and evaluating large-scale systems.
Featured Work
Language Model Benchmark Contamination Leaves a Person-Fit Signature
Under review — AAAI-27
A benchmark score is trustworthy only if the evaluated model has not trained on the test items, yet existing contamination detectors demand privileged access — model weights, token log-probabilities, or the training corpus — that a public leaderboard entry does not expose. We propose a detector requiring only the graded response matrix: which items each model answered correctly. Treating contamination as psychometric item preknowledge, a contaminated model behaves like an aberrant test-taker, succeeding on items too difficult for its true ability because it has memorized them.
See my research & projects or
read my CV. You can reach me at
aveizi (at) stanford (dot) edu, or find me on
LinkedIn and
GitHub.