Чек-лист · Digital Economy Lab Checklist · Digital Economy Lab

Как оценить LLM для банка

How to evaluate an LLM for a bank

10 вопросов, на которые команде стоит ответить до вывода LLM-агента в продуктив. Отметьте то, что у вас уже есть, — страница посчитает итог. Ничего никуда не отправляется.

10 questions a team should answer before putting an LLM agent into production. Tick what you already have and the page adds it up. Nothing is sent anywhere.

Зачем

Why

LLM-агента в банке мало запустить — нужно доказать при проверке, что ему можно доверять. Демо «работает хорошо» не отвечает на главный вопрос риск-функции и регулятора: что будет с качеством, когда меняются правила, данные или сама модель.

Launching an LLM agent in a bank is not enough — you have to show an examiner that it can be trusted. A demo that "works well" does not answer the question risk and supervisors actually ask: what happens to quality when the rules, the data or the model itself change.

Каждый «нет» ниже — риск, который всплывёт на валидации модели или внешней проверке.

Every "no" below is a risk that will surface at model validation or an external review.

Чек-лист: 10 вопросов

Checklist: 10 questions

01Уровень диалога — ответ клиенту или сотрудникуDialogue level — the answer to a customer or employee
02Уровень задачи — бизнес-процессTask level — the business process
03Уровень агента — поведение во времениAgent level — behaviour over time
0 / 10

Как проверить это на практике

How to test it in practice

CACE-Bench — открытый синтетический бенчмарк Digital Economy Lab для LLM-агентов кредитного конвейера. Он моделирует смену трактовки комплаенс-нормы и измеряет, восстанавливается ли агент без регрессии на ранее верных кейсах. Внутри: синтетические заявки и реестр организаций (без персональных данных), размеченные трассы агентов и метрики трёх уровней из чек-листа.

CACE-Bench is Digital Economy Lab's open synthetic benchmark for LLM credit-pipeline agents. It simulates a change in how a compliance rule is interpreted and measures whether the agent recovers without regressing on previously correct cases. Inside: synthetic applications and an organisation registry (no personal data), labelled agent traces and the three metric tiers from this checklist.

Что мы предлагаем командам риска, комплаенса и Data/ML

What we offer risk, compliance and data/ML teams

Вебинар и демоWebinar and demoРазбор CACE-Bench на примере кредитного конвейера — бесплатно.A walkthrough of CACE-Bench on a credit pipeline — free.
Воркшоп, 1 деньWorkshop, 1 dayКарта метрик под ваш процесс.A metric map for your own process.
Программа, 4–6 недельProgramme, 4–6 weeksМетрики, LLM-судья, наблюдаемость, аудит-трейл и итоговый план аудита.Metrics, LLM judge, observability, audit trail and a final audit plan.

Равиль Ахтямов, digital-экономист · Digital Economy Lab

Ravil Akhtyamov, digital economist · Digital Economy Lab

Made on
Tilda