BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data Paper • 2510.10159 • Published Oct 11, 2025 • 3
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation Paper • 2605.22544 • Published May 21 • 2