Most models can sound profound. Few can hold a precise position when you argue back.
A closed-book eval of classical Advaita Vedānta: school boundaries, levels of reality, text-grounded reading, and multi-turn pressure. Built to surface confident vagueness and doctrine that flips — the reliability failures that trivia benchmarks miss.
The headline score, AdvaitaBench-N, averages two AI judges after correcting each judge's measured bias toward its own maker's models. 50 is the field average. Hover any point for detail.
Each line is one model. The left column is an Anthropic judge with a permissive rubric, the right an OpenAI judge with a strict one. Lines that cross tell you the ranking depends on who grades, which is exactly why the headline score uses both.
Three design decisions do most of the work. Full methodology, rubrics, and the bias-correction formula are in the blog and the repo.
All 84 tasks are public. Filter by family. "Novel" tasks were written for this benchmark and exist nowhere in training data; "multi-turn" tasks include scripted pushback.