← Digital Economy Lab · все материалы← Digital Economy Lab · all work
Бенчмарк, который не закрывается масштабом. Задачи хорезмской математической школы — аль-Хорезми, аль-Бируни, аль-Каши — которые языковая модель должна решить не только правильно, но и методом самого источника.A benchmark that scale does not close. Problems from the Khwarezm mathematical school — al-Khwarizmi, al-Biruni, al-Kashi — which a language model has to solve not only correctly but by the source's own method.
«Какой квадрат, сложенный с десятью своими корнями, даёт тридцать девять?» — то есть x² + 10x = 39. Правильный ответ у обеих моделей ниже одинаковый: x = 3. Различается путь.“What must be the square which, when increased by ten of its own roots, amounts to thirty-nine?” — that is, x² + 10x = 39. Both models below reach the right answer, x = 3. The path differs.
x = (−10 + √(100 + 156)) / 2 = 3. Верно, но этого метода у аль-Хорезми нет: это анахронизм.x = (−10 + √(100 + 156)) / 2 = 3. Correct, but al-Khwarizmi has no such method: it is an anachronism.
Геометрия за этими шагами — она же логотип бенчмарка. Светлый квадрат в центре — x². Десять корней разложены на четыре полосы по 2,5x. Четыре пунктирных угла по 2,5 × 2,5 добавляют 25 и дополняют фигуру до квадрата площадью 39 + 25 = 64 со стороной 8, откуда x = 8 − 5 = 3.The geometry behind those steps — also the benchmark's logo. The light square in the centre is x². The ten roots are laid out as four strips of 2.5x. Four dashed corners of 2.5 × 2.5 add 25 and complete the figure to a square of area 39 + 25 = 64 with side 8, whence x = 8 − 5 = 3.
Решает ли языковая модель задачи из «Китаб аль-джабр» аль-Хорезми и родственных хорезмских источников методом самого источника. Правильность ответа и методологическая верность оцениваются раздельно; верность размечают вслепую люди по рубрике из девяти критериев, без LLM-судьи. В верность входит следование исторической таксономии шести случаев и геометрии дополнения до квадрата — измерение, которого нет в стандартных математических бенчмарках.Whether a language model solves problems from al-Khwarizmi's Kitab al-jabr and companion Khwarezm-school sources by the source's own method. Correctness and methodological fidelity are scored separately; fidelity is rated by blind human annotators on a nine-criterion rubric, with no LLM judge. Fidelity covers adherence to the historical six-case taxonomy and completing-the-square geometry — a dimension absent from standard mathematical benchmarks.
Выбор корпуса фронтирной лабораторией смоделирован как задача распределения под бюджетным ограничением. При законе масштабирования, обусловленном носителем, и допущении незаменимости разрыв на бенчмарке имеет нижнюю границу δ > 0, равномерную по вычислениям: ни один конечный множитель compute его не закрывает. Жёсткая граница — корпуса нет в машиночитаемом виде — отделена от мягкой, где он есть, но приобретать его невыгодно; это отделение объясняет, почему открытая выкладка не разрушает утверждение.The frontier laboratory's corpus choice is modelled as a budget-constrained allocation problem. Under a support-conditioned scaling law and a non-substitutability assumption, the benchmark gap admits a lower bound δ > 0 that is uniform in compute: no finite compute multiplier closes it. A hard bound — the corpus does not exist in machine-readable form — is separated from a soft one, where it exists but acquiring it is never optimal; that separation is why an open release does not dissolve the claim.
Это пилот: выложены 15 позиций из 100 запланированных. Заявленная корреляция «известность — верность» требует не меньше 33 позиций; пилот из 15 даёт 43 % мощности, и мы сообщаем мощность, а не результат, которого дизайн не выдерживает. Многомодельное исследование верности с человеческой разметкой в работе. Трёхъязычный выровненный корпус (арабский — русский — узбекский) готовится; в версии 0.1.0 опубликована его схема. Текст современных изданий под копирайтом не воспроизводится.This is a pilot: 15 of the 100 planned items are released. The fame–fidelity correlation claim requires at least 33 items; a pilot of 15 attains 43% power, and we report the power rather than a result the design cannot support. The multi-model, human-scored fidelity study is in progress. The tri-lingual aligned corpus (Arabic — Russian — Uzbek) is in preparation; version 0.1.0 releases its schema. No text of copyrighted modern editions is reproduced.
| data/items.json | 15 открытых позиций пилота: условие, эталонный ответ и вывод, источник, уровень известности15 public pilot items: statement, gold answer and derivation, source, fame level |
| rubric/RUBRIC.md | Рубрика верности из девяти критериев и формула оценкиThe nine-criterion fidelity rubric and the scoring formula |
| harness/ | Сбор ответов моделей, слепые перемешанные листы для разметчиков, согласие и корреляцияCollecting model outputs, blind shuffled annotation sheets, agreement and correlation |
| stats/ | Скрипт, воспроизводящий все числа анализа мощности из статьиA script reproducing every number of the paper's power analysis |
Как цитировать: Akhtyamov R. Khwarezm-100: an access-bounded benchmark of medieval Khwarezm-school mathematics. Версия 0.1.0. Digital Economy Lab, 2026. DOI 10.5281/zenodo.23057573.How to cite: Akhtyamov R. Khwarezm-100: an access-bounded benchmark of medieval Khwarezm-school mathematics. Version 0.1.0. Digital Economy Lab, 2026. DOI 10.5281/zenodo.23057573.
← Digital Economy Lab · все материалы← Digital Economy Lab · all work