The Embedder's Dilemma: LLMs Are Better, but at What Cost?
Paper • 2608.12875 • Published • 14
None defined yet.
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models