MarkTechPost수정일

Laya 개발자 가이드: 제로샷 의사결정과 캘리브레이션

오픈소스 제로샷 의사결정 엔진인 Laya에 대한 포괄적인 코딩 가이드를 살펴보세요. 타입화된 결정(typed decisions) 구현, 커스텀 온도 피팅, 실제 CLINC150 금융 데이터를 활용한 신뢰할 수 있는 거절 게이트(abstention gate) 구축 방법을 배워보세요. 이 글 Laya 개발자 가이드: 제로샷 의사결정과 캘리브레이션은 MarkTechPost 에서 처음…

이 튜토리얼에서는 다음을 사용합니다. LayaConvai Innovations가 공개한 오픈소스 의사결정 엔진으로, 2026년 9월 가장 많은 스타를 받은 머신러닝 저장소 중 하나가 된 프로젝트입니다. Laya는 비자기회귀(non-autoregressive) System 1 모델입니다. 텍스트를 생성하는 대신, 421-million-parameter 인코더가 텍스트 한 조각과 유형화된 질문들—레이블 중 선택, 척도 위의 점수, 또는 예/아니오—을 읽고, 출력 토큰을 전혀 생성하지 않으면서 단일 forward pass로 모든 옵션에 대한 확률을 반환합니다. 이 모델의 강점은 속도와 보정된 확률(calibrated probabilities)이며, TypeSafe의 Jev에 대한 열린 대안입니다. README의 예제를 반복하는 대신, 우리는 이러한 약속을 정답이 알려진 실제 레이블 데이터, 즉 CLINC150 의도 데이터셋의 은행 도메인에 적용하여 프로덕션 라우터가 실제로 얻는 것을 측정했습니다: 학습된 분류기 대비 zero-shot 정확도, 옵션의 표현과 순서가 얼마나 중요한지, 제공되는 확률이 얼마나 정직한지, 검증 데이터에 온도(temperature)를 피팅하면 무엇이 해결되고 무엇이 조용히 깨지는지, 오류 예산에 맞춘 기권(abstention) 게이트, 범위 외(out-of-scope) 트래픽, 온도로는 복구할 수 없는 예/아니오 질문, 그리고 pydantic 스키마에서 나오는 유형화된 출력입니다.

코드 복사
import os
import sys
import time
import json
import warnings
import traceback
import subprocess
import urllib.request
 
RESULTS = {}
 
 
def banner(title):
    print("\n" + "=" * 78)
    print(title)
    print("=" * 78)
 
 
def section(name):
    def wrap(fn):
        def run(*a, **kw):
            banner(name)
            try:
                out = fn(*a, **kw)
                RESULTS[name] = out if isinstance(out, str) else "ok"
                return out
            except Exception as e:
                RESULTS[name] = f"SKIPPED / FAILED -> {type(e).__name__}: {e}"
                print(f"\n[!] {name} did not complete: {type(e).__name__}: {e}")
                traceback.print_exc(limit=3)
                return None
        return run
    return wrap
 
 
banner("1. Install Laya and load the English checkpoint at its reviewed revision")
subprocess.run([sys.executable, "-m", "pip", "install", "-q", "laya==0.3.27"], check=True)
 
import numpy as np
import pandas as pd
import torch
import laya
from laya.calibrate import records_from_labeled
from laya.evals import selective_accuracy, aurc
from laya.common import temp_bucket
 
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
# laya.load() follows the Hub's main branch unless told otherwise. The package ships the commit
# its authors reviewed for each checkpoint; pinning it keeps this notebook's weights fixed.
REVISION = laya.PINNED_REVISIONS["convaiinnovations/laya"]
with warnings.catch_warnings(record=True) as caught:
    warnings.simplefilter("always")
    agent = laya.load("convaiinnovations/laya", device=DEVICE, revision=REVISION)
# On CUDA Laya autocasts to fp16/bf16. Turning that off keeps every device in fp32, so a GPU run
# reproduces the CPU numbers printed below to within floating-point noise.
agent.amp_enabled = False
SHIPPED = (list(agent.temperature), dict(agent.temperature_by_options))
 
n_params = sum(p.numel() for p in agent.model.parameters())
print(f"  laya {laya.__version__}  |  torch {torch.__version__}  |  device {DEVICE}, fp32")
print(f"  checkpoint convaiinnovations/laya @ {REVISION[:7]}  |  {n_params / 1e6:.0f}M parameters"
      f"  |  max_len {agent.cfg['max_len']}, head_max_len {agent.cfg['head_max_len']}")
print("\n  Temperatures shipped with the checkpoint (probabilities = softmax(logits / T)):")
for qt, name in enumerate(["choice", "score", "noul"]):
    print(f"    {name:7s} type-level T = {SHIPPED[0][qt]:.3f}")
for bucket, t in sorted(SHIPPED[1].items()):
    print(f"    {bucket:12s} T = {t:.3f}")
for w in caught:
    if "temperature" in str(w.message):
        print("\n  Warning at load time:\n    " + str(w.message).replace("; ", ";\n    "))
print("\n  T > 1 softens probabilities and T < 1 sharpens them. Remember the choice:11+ row: the")
print("  checkpoint ships 0.10 there, which the loader clamps to 0.5, so any choice question with")
print("  11 or more options gets probabilities SHARPENED by 2x. Step 6 measures what that costs.")

릴리스된 패키지인 laya 0.3.27을 설치하고 영어 체크포인트를 로드합니다. 여기서 두 가지 선택이 실행을 재현 가능하게 만듭니다. 기본적으로 laya.load는 Hugging Face main 브랜치를 따르므로, 라이브러리 저자들이 검토한 리비전, 즉 laya.PINNED_REVISIONS로 노출되는 리비전을 고정합니다. 그리고 CUDA에서 Laya는 자동으로 half precision으로 변환되므로, 이를 꺼서 모든 디바이스를 fp32로 유지하고 GPU 실행이 여기에 표시된 CPU 수치를 재현하도록 합니다. 체크포인트에 포함된 온도를 출력해 보면 예측을 수행하기 전에 첫 번째 발견이 나옵니다: 옵션이 11개 이상인 choice 질문의 항목이 유효 범위를 벗어난 0.10으로, 로더가 이를 0.5로 클램프하고 경고합니다. 1 미만의 온도는 확률을 날카롭게 만드므로, 그만큼의 옵션을 가진 질문에 대한 모든 답변은 원래 모델보다 두 배 더 확실해 보일 것입니다.

코드 복사
TICKET = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
TRIAGE = {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
                   "criteria": {"billing": "invoices, payments, refunds",
                                "technical": "bugs, outages, system errors",
                                "other": "everything else"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["not urgent", "soon", "blocking"]},
    "churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},
}
 
 
@section("2. One forward pass, three typed questions, zero output tokens")
def first_decision():
    r = agent.predict(TICKET, TRIAGE)
    a = r["answers"]
    d, u, c = a["department"], a["urgency"], a["churn_risk"]
    print(f"  state: {TICKET!r}\n")
    print(f"  department (choice)  -> {d['choice']!r}   probabilities {d['probabilities']}")
    print(f"  urgency    (score)   -> {u['score']:.2f} on 0..2   level probabilities {u['probabilities']}")
    print(f"  churn_risk (noul)    -> P(yes) = {c['noul']:.3f}")
    print("\n  Two confidence fields, two different quantities:")
    for qid, ans in a.items():
        print(f"    {qid:11s} answer_confidence {ans['answer_confidence']:.3f}   confidence {ans['confidence']:.3f}")
    print("  answer_confidence is the probability of the reported answer, max(p): the number that")
    print("  calibration, the abstention gate and every metric below use. confidence is 1 - normalised")
    print("  entropy, whose scale depends on the number of options. Do not threshold on it.")
    print(f"\n  usage: {r['usage']}")
    print("  output_tokens is always 0: Laya scores the options it is given and never generates text.")
    return f"{d['choice']} / urgency {u['score']:.2f} / P(churn) {c['noul']:.2f} in one pass"
 
 
first_decision()

predict를 한 번 호출하면 지원 티켓에 대한 세 개의 타입화된 질문을 단일 forward pass로 답변합니다: 부서는 choice로, 긴급도는 0부터 2까지의 점수로, 이탈 위험은 예/아니오로 답변합니다. 결과에는 모든 옵션에 대한 확률과 혼동하기 쉬운 두 개의 신뢰도 필드가 담겨 있습니다. answer_confidence는 보고된 답변의 확률이며, 이 튜토리얼 뒤에서 캘리브레이션, 거절 게이트 및 모든 지표가 사용하는 값입니다. confidence는 1에서 정규화된 엔트로피를 뺀 값으로, 그 스케일은 질문의 옵션 수에 따라 달라집니다. usage 블록은 출력 토큰이 0임을 보여주는데, Laya는 주어진 옵션을 스코어링할 뿐 텍스트를 생성하지 않기 때문입니다.

코드 복사
@section("3. What a pass costs: questions are rows, options are nearly free")
def cost_model():
    def median_ms(q, n=7):
        agent.predict(TICKET, q)
        times = []
        for _ in range(n):
            t0 = time.perf_counter()
            r = agent.predict(TICKET, q)
            times.append(1000 * (time.perf_counter() - t0))
        return float(np.median(times)), r["usage"]["input_tokens"]
 
    print(f"  {'one state, asking ...':34s} {'ms (median of 7)':>16s} {'input_tokens':>13s}")
    rows = {}
    for n in (1, 4, 16):
        q = {f"q{i}": {"type": "noul", "instructions": f"Does the message mention topic number {i}?"} for i in range(n)}
        rows[f"{n} yes/no"] = median_ms(q)
        print(f"  {f'{n:2d} yes/no questions':34s} {rows[f'{n} yes/no'][0]:16.1f} {rows[f'{n} yes/no'][1]:13d}")
    for k in (3, 15, 40):
        q = {"pick": {"type": "choice", "instructions": "Pick one", "criteria": [f"option {i}" for i in range(k)]}}
        rows[f"{k} options"] = median_ms(q)
        print(f"  {f'1 choice question, {k:2d} options':34s} {rows[f'{k} options'][0]:16.1f} {rows[f'{k} options'][1]:13d}")
    print("\n  Every question is encoded as its own (state, question) row, so 16 yes/no questions cost")
    print("  roughly 16 rows. All options of one choice question share a single row and its head")
    print("  budget, so 40 options cost far less than 40 yes/no questions. Design rule: ask one")
    print("  choice question with many options, not many yes/no questions.")
    return (f"16 yes/no {rows['16 yes/no'][0]:.0f} ms vs one 40-option choice "
            f"{rows['40 options'][0]:.0f} ms")
 
 
cost_model()

Laya 위에 무언가를 만들기 전에, 단일 메시지에 대한 forward pass의 비용을 측정합니다. 각 질문은 메시지와 쌍을 이루는 자체 행이 되므로, 16개의 예/아니오 질문은 하나보다 약 8배 더 오래 걸립니다. choice 질문의 모든 옵션은 하나의 행과 그 헤드 예산을 공유하므로, 40개 옵션 choice는 3개 옵션 질문보다 겨우 2배이며, 우리 CPU에서 16개의 예/아니오 질문 비용의 4분의 1에 불과합니다. 이는 이후 모든 것을 형성하는 설계 규칙을 제공합니다: 많은 예/아니오 질문보다 많은 옵션을 가진 하나의 choice 질문을 하라는 것입니다.

코드 복사
CARD = json.load(urllib.request.urlopen("https://huggingface.co/api/datasets/clinc/clinc_oos"))["cardData"]
PLUS = next(c for c in CARD["dataset_info"] if c["config_name"] == "plus")
NAMES = {int(k): v for k, v in PLUS["features"][1]["dtype"]["class_label"]["names"].items()}
 
 
def clinc(split):
    df = pd.read_parquet(f"https://huggingface.co/api/datasets/clinc/clinc_oos/parquet/plus/{split}/0.parquet")
    return df.assign(label=df.intent.map(NAMES))
 
 
# The banking domain of CLINC150 (15 intents), with the one-line descriptions a developer would write.
DESCRIBED = {
    "balance": "checking how much money is in an account",
    "transactions": "looking up recent transactions on an account",
    "transfer": "moving money between accounts or to another person",
    "freeze_account": "freezing or locking an account",
    "account_blocked": "an account that is blocked or locked and cannot be used",
    "pay_bill": "paying a bill",
    "bill_balance": "how much is owed on a bill",
    "bill_due": "when a bill is due",
    "interest_rate": "the interest rate on an account",
    "min_payment": "the minimum payment that is due",
    "order_checks": "ordering new checks or a checkbook",
    "pin_change": "changing a PIN",
    "report_fraud": "reporting fraud or suspicious activity",
    "routing": "the bank routing number",
    "spending_history": "how much was spent over a period or on a category",
}
INTENTS = list(DESCRIBED)
ASK = "Which banking request is this?"
 
 
def route(states, criteria, **kw):
    """One choice question over `criteria` for every state; returns choices, confidences, results."""
    res = agent.predict_batch(list(states), {"intent": {"type": "choice", "instructions": ASK, "criteria": criteria}},
                              batch_size=32, **kw)
    return (np.array([r["answers"]["intent"]["choice"] for r in res]),
            np.array([r["answers"]["intent"]["answer_confidence"] for r in res]), res)
 
 
@section("4. Real labelled data: CLINC150 banking, zero-shot vs. a trained classifier")
def zero_shot_vs_trained():
    global train, val, test, BTRAIN, BVAL, BTEST
    train, val, test = clinc("train"), clinc("validation"), clinc("test")
    BTRAIN, BVAL, BTEST = (d[d.label.isin(INTENTS)].reset_index(drop=True) for d in (train, val, test))
    print(f"  CLINC150 'plus': {len(train):,} / {len(val):,} / {len(test):,} train / validation / test queries,"
          f" 150 intents + out-of-scope")
    print(f"  banking domain: {len(INTENTS)} intents; {len(BTRAIN)} train, {len(BVAL)} validation, {len(BTEST)} test queries")
    print(f"  e.g. {BTEST.text[0]!r} -> {BTEST.label[0]}")
 
    t0 = time.time()
    pred, conf, _ = route(BTEST.text, DESCRIBED)
    secs = time.time() - t0
    acc = float((pred == BTEST.label).mean())
    globals()["DESCRIBED_ACC"], globals()["DESCRIBED_PRED"] = acc, pred
    print(f"\n  Laya, zero-shot, 15 options with descriptions: accuracy {acc:.3f}"
          f"   ({secs:.0f}s for {len(BTEST)} queries on {DEVICE})")
 
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression
    print("\n  TF-IDF + logistic regression trained on k labelled queries per intent (5 draws for k < 100):")
    curve = {}
    for k in (1, 3, 10, 30, 100):
        accs = []
        for seed in range(5 if k < 100 else 1):
            sub = BTRAIN.groupby("label", group_keys=False).sample(k, random_state=seed)
            vec = TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True)
            clf = LogisticRegression(max_iter=3000, C=10).fit(vec.fit_transform(sub.text), sub.label)
            accs.append(float((clf.predict(vec.transform(BTEST.text)) == BTEST.label).mean()))
        curve[k] = float(np.mean(accs))
        print(f"    k = {k:3d}  ({k * len(INTENTS):5,d} labels)   accuracy {curve[k]:.3f}  (sd {np.std(accs):.3f})")
    globals()["CURVE"] = curve
    print("\n  Zero-shot Laya, with nothing but the intent descriptions, lands between what a classic")
    print("  classifier reaches with three and with ten labelled examples per intent. Step 5 shows the")
    print("  descriptions are the weak part.")
    return f"zero-shot {acc:.3f}; TF-IDF needs 10/intent for {curve[10]:.3f}"
 
 
zero_shot_vs_trained()

실제 레이블된 데이터를 위해 우리는 CLINC150을 사용합니다. 이는 10개 도메인에 걸친 150개 인텐트와 범위 외 질의 집합으로 구성된 공개 인텐트 분류 벤치마크로, parquet 파일로 Hugging Face Hub에서 직접 읽어옵니다. 우리는 금융 도메인, 즉 인텐트당 학습 100개, 검증 20개, 테스트 30개의 질의를 가진 15개 인텐트를 사용하고, Laya가 450개 테스트 질의를 제로샷으로 라우팅하도록 하며, 각 인텐트의 이름과 개발자가 작성할 법한 한 줄 설명을 제공합니다. 레이블된 예제 없이 0.804의 정확도에 도달합니다. 규모 비교를 위해, TF-IDF 및 로지스틱 회귀 분류기는 인텐트당 레이블된 질의 3개로 0.651, 10개로 0.848, 30개로 0.904에 도달합니다.

코드 복사
@section("5. Criteria wording and option order change the answers")
def wording_and_order():
    t0 = time.time()
    pred_n, conf_n, res_n = route(BTEST.text, INTENTS)
    secs = time.time() - t0
    acc_n = float((pred_n == BTEST.label).mean())
    pred_r, _, _ = route(BTEST.text, INTENTS[::-1])
    acc_r = float((pred_r == BTEST.label).mean())
    flips = float((pred_r != pred_n).mean())
    print(f"  {'criteria':44s} {'accuracy':>8s}")
    print(f"  {'15 names with one-line descriptions (step 4)':44s} {DESCRIBED_ACC:8.3f}")
    tok = {name: agent.predict(BTEST.text[0], {"intent": {"type": "choice", "instructions": ASK, "criteria": c}})
           ["usage"]["input_tokens"] for name, c in (("described", DESCRIBED), ("names", INTENTS))}
    print(f"  {'15 bare intent names':44s} {acc_n:8.3f}   ({secs:.0f}s; {tok['names']} input tokens per query"
          f" vs {tok['described']})")
    print(f"  {'15 bare intent names, order reversed':44s} {acc_r:8.3f}")
    print(f"\n  Reversing the option order changes {flips:.1%} of individual answers, even where the")
    print("  overall accuracy barely moves: the model has a position prior, so fix the order you deploy.")
    changed = pd.DataFrame({"label": BTEST.label, "described": DESCRIBED_PRED, "names": pred_n})
    gained = changed[(changed.names == changed.label) & (changed.described != changed.label)]
    lost = changed[(changed.names != changed.label) & (changed.described == changed.label)]
    print(f"\n  Bare names fixed {len(gained)} answers the descriptions got wrong and broke {len(lost)};"
          f" the most common fixes:")
    for (lab, was), n in gained.groupby(["label", "described"]).size().sort_values(ascending=False).head(3).items():
        print(f"    {lab:16s} had been routed to {was:16s} x{n}")
    print("\n  More text is not more signal. The descriptions we wrote blurred intents the names keep")
    print("  apart, and only labelled data could tell us. From here on we route on the bare names.")
    globals().update(PRED=pred_n, CONF=conf_n)
    return f"descriptions {DESCRIBED_ACC:.3f} -> names {acc_n:.3f}; reversed order flips {flips:.1%}"
 
 
wording_and_order()

다음으로 옵션의 표현만 바꿉니다. 우리의 설명 없이 15개의 베어(bare) 인텐트 이름만을 Laya에 제공하면, 옵션이 절반 이하의 토큰을 차지하기 때문에 정확도가 0.804에서 0.878로 상승하고 시간이 절반으로 줄어듭니다. 설명은 이름들이 구분해 주는 인텐트들을 흐릿하게 만들었습니다: account_blocked가 freeze_account로 10번 라우팅되었고, 이자율 질문은 balance로 라우팅되었습니다. 베어 이름의 순서를 뒤집으면 전체 정확도는 거의 움직이지 않음에도 개별 답변의 4.2 퍼센트가 바뀌는데, 이는 위치 사전(position prior)의 징후이므로, 배포하는 옵션 순서는 테스트한 순서여야 합니다. 이 둘 중 무엇이든 알려줄 수 있는 것은 레이블된 데이터뿐입니다. 이후부터는 베어 이름으로 라우팅합니다.

Copy Code
def reliability(conf, correct, title):
    bins = np.linspace(0, 1, 11)
    idx = np.clip(np.digitize(conf, bins) - 1, 0, 9)
    print(f"  {title}\n    {'confidence':>12s} {'answers':>8s} {'mean conf':>10s} {'accuracy':>9s}")
    for i in range(10):
        m = idx == i
        if m.sum():
            print(f"    {bins[i]:5.1f}-{bins[i + 1]:.1f}   {m.sum():8d} {conf[m].mean():10.3f} {correct[m].mean():9.3f}")
 
 
@section("6. How honest are the probabilities as shipped?")
def shipped_calibration():
    ok = PRED == BTEST.label.values
    ece = laya.ece_score(CONF, ok)
    bucket = temp_bucket(laya.QTYPES["choice"], len(INTENTS))
    print(f"  15-option question -> temperature bucket {bucket!r}, T = {agent.temperature_by_options.get(bucket):.2f}\n")
    reliability(CONF, ok, "Reliability on the 450 test queries:")
    print(f"\n  accuracy {ok.mean():.3f}, mean answer_confidence {CONF.mean():.3f}, ECE {ece:.3f} (15 bins)")
    top = CONF >= 0.9
    print(f"  {top.mean():.0%} of answers claim >= 0.9 confidence; {ok[top].mean():.1%} of those are right.")
    print("\n  The model is over-confident here, and the clamp is part of why: with T = 0.5 every")
    print("  15-option answer is sharpened before you see it. 'Calibrated' in a model card describes a")
    print("  training objective, not a property of your question. Measure it on your own labels.")
    globals()["ECE_SHIPPED"] = ece
    return f"ECE {ece:.3f} as shipped; accuracy {ok.mean():.3f} at mean confidence {CONF.mean():.3f}"
 
 
shipped_calibration()

이제 확률이 얼마나 정직한지 묻습니다. 15개 선택지 문제는 체크포인트의 choice:11+ temperature 버킷, 즉 0.5로 고정된 버킷에 해당합니다. 450개 테스트 쿼리에 대한 신뢰도 표는 그 결과를 보여줍니다: 답변의 92 퍼센트가 최소 0.9의 신뢰도를 주장하지만, 그중 91.1 퍼센트만 옳으며, 0.974의 평균 신뢰도는 0.878의 정확도보다 훨씬 높아 기대 보정 오차(expected calibration error)는 0.102입니다. Laya의 학습 목적함수는 적절한 스코어링 규칙(proper scoring rules)을 사용하며, 이것이 모델 카드에서 보정됨(calibrated)이 의미하는 바입니다. 그럼에도 불구하고 보정은 어떤 분포 위의 질문에 대한 성질이므로, 여러분 자신의 레이블로 측정해야 합니다.

Copy Code
@section("7. Fit a temperature on validation data, without breaking the other questions")
def fit_temperature():
    eye = np.eye(len(INTENTS))
    pairs = [(s, {"intent": {"type": "choice", "instructions": ASK, "criteria": INTENTS}},
              {"intent": eye[INTENTS.index(l)]}) for s, l in zip(BVAL.text, BVAL.label)]
    t0 = time.time()
    records = records_from_labeled(agent, pairs)
    print(f"  {len(records)} labelled validation records (raw logits + one-hot targets) in {time.time() - t0:.0f}s")
 
    fit = agent.fit_temperatures(records)
    after = (list(agent.temperature), dict(agent.temperature_by_options))
    print(f"  fitted on bucket counts {fit['n_by_bucket']}; choice temperature {fit['temperature'][0]:.3f}")
    print("\n  What agent.fit_temperatures() just installed, next to what shipped:")
    print(f"    {'':14s} {'shipped':>8s} {'after':>8s}")
    for qt, name in enumerate(["choice", "score", "noul"]):
        print(f"    {name + ' (type)':14s} {SHIPPED[0][qt]:8.3f} {after[0][qt]:8.3f}")
    for bucket in sorted(SHIPPED[1]):
        now = after[1].get(bucket)
        print(f"    {bucket:14s} {SHIPPED[1][bucket]:8.3f} {('%.3f' % now) if now is not None else '  (gone)':>8s}")
    print("\n  One fit on choice questions replaced the whole map: every per-bucket entry is gone (a bucket")
    print("  needs 2,000 records to keep its own temperature) and the score and noul temperatures were reset")
    print("  to 1.0 because there were no records of those types. Every yes/no question now uses a")
    print("  different temperature than it did a minute ago. Install only what you measured instead:")
    agent.temperature, agent.temperature_by_options = list(SHIPPED[0]), dict(SHIPPED[1])
    agent.temperature_by_options["choice:11+"] = fit["temperature"][0]
    print(f"    agent.temperature_by_options['choice:11+'] = {fit['temperature'][0]:.3f}   (all else as shipped)")
 
    pred, conf, _ = route(BTEST.text, INTENTS)
    ok = pred == BTEST.label.values
    ece = laya.ece_score(conf, ok)
    print()
    reliability(conf, ok, "Reliability on the same 450 test queries, after the fit:")
    print(f"\n  ECE {ECE_SHIPPED:.3f} -> {ece:.3f}; accuracy unchanged at {ok.mean():.3f} (temperature never changes the argmax)")
    agent.save_calibration("laya_banking_calibration.json")
    saved = json.load(open("laya_banking_calibration.json"))
    print(f"  saved with agent.save_calibration(): keys {sorted(saved)}; reload with laya.load(..., calibration=path)")
    globals().update(RECORDS=records, CONF_FIT=conf, OK_FIT=ok)
    return f"ECE {ECE_SHIPPED:.3f} -> {ece:.3f} from {len(records)} validation queries"
 
 
fit_temperature()

Laya의 보정 모듈은 레이블이 달린 예시를 원시 로짓(raw logits)과 타깃의 기록으로 변환하고, 그에 맞게 temperature를 피팅합니다. 우리는 검증 분할에서 300개의 기록을 만들고 agent.fit_temperatures를 호출하는데, 이는 1.258의 choice temperature를 피팅하여 설치한 뒤, 출하된 값 옆에 전체 temperature 표를 출력합니다. 이 피팅은 맵 전체를 교체했습니다: 옵션 수별 항목은 모두 사라졌는데, 버킷이 자체 temperature를 유지하려면 2,000개의 기록이 필요하기 때문이며, score와 yes/no temperature는 해당 유형의 기록이 없어 1.0으로 초기화되었습니다. choice 문제에 대한 한 번의 피팅이 에이전트의 모든 yes/no 질문이 보정되는 방식을 조용히 바꿔버린 것입니다. 그래서 우리는 출하된 값을 복원하고 측정한 버킷만 설치합니다. 테스트 세트에서 보정 오차는 정확도가 변하지 않은 채 0.102에서 0.059로 떨어지는데, temperature는 어느 옵션이 이기는지를 절대 바꾸지 않기 때문입니다. 그리고 save_calibration이 결과를 laya.load가 다시 읽을 수 있는 JSON 파일에 기록합니다.

Copy Code
@section("8. An abstention gate fitted to an error budget")
def abstention_gate():
    t = agent.temperature_by_options["choice:11+"]
    z = np.array([r[1] for r in RECORDS]) / t
    p_val = np.exp(z - z.max(1, keepdims=True)); p_val /= p_val.sum(1, keepdims=True)
    conf_val = p_val.max(1)
    ok_val = np.array([np.argmax(r[1]) == np.argmax(r[2]) for r in RECORDS])
    print(f"  {'target':>6s} {'threshold':>10s}   {'validation: kept / error':>25s}   {'test: kept / error':>19s}")
    gates = {}
    for target in (0.02, 0.05, 0.10):
        thr = laya.fit_abstention_thresholds(RECORDS, agent.temperature, agent.temperature_by_options,
                                             target_error=target)
        cut = thr["choice:11+"]
        kv, kt = conf_val >= cut, CONF_FIT >= cut
        gates[target] = thr
        print(f"  {target:6.0%} {cut:10.3f}   {kv.mean():14.1%} / {1 - ok_val[kv].mean():5.1%}"
              f"   {kt.mean():11.1%} / {1 - OK_FIT[kt].mean():5.1%}")
    _, _, res = route(BTEST.text, INTENTS, min_confidence=gates[0.05])
    states = pd.Series([r["answers"]["intent"]["abstention"] for r in res]).value_counts().to_dict()
    print(f"\n  The 5% gate applied by Laya itself, predict_batch(min_confidence=...): {states}")
    conf, ok = CONF_FIT.tolist(), OK_FIT.tolist()
    print(f"  selective accuracy on test: top 50% by confidence {selective_accuracy(conf, ok, 0.5):.3f},"
          f" top 80% {selective_accuracy(conf, ok, 0.8):.3f}, all {np.mean(ok):.3f};  AURC {aurc(conf, ok):.3f}")
    print(f"\n  On validation every cut meets its target, by construction. On test the realised error is")
    print(f"  higher, because the test queries are harder: accuracy is {ok_val.mean():.3f} on validation and"
          f" {OK_FIT.mean():.3f}")
    print("  on test. A gate fitted to an error budget holds only for traffic that looks like the data it")
    print("  was fitted on. Fit it on a sample of real traffic, re-check it as traffic drifts, and leave")
    print("  margin. Thresholds are per option-count bucket because one number does not transfer between")
    print("  a 2-way and a 15-way question.")
    globals()["GATE"] = gates[0.05]
    kept = CONF_FIT >= gates[0.05]["choice:11+"]
    return f"5% target: {1 - OK_FIT[kept].mean():.1%} error on test at {kept.mean():.1%} coverage"
 
 
abstention_gate()

laya.fit_abstention_thresholds는 동일한 검증 기록을 받아, 각 옵션 수 버킷에 대해 검증 오차를 목표 안에 유지하는 가장 느슨한 신뢰도 컷을 반환합니다. 5 퍼센트 목표에 대해 0.602를 선택하는데, 이는 검증 쿼리의 95.7 퍼센트를 4.5 퍼센트 오차로 유지하며, Laya는 이 값이 min_confidence로 predict_batch에 전달될 때 스스로 동일한 컷을 적용하여 450개 테스트 답변 중 35개를 기권(abstained)으로 표시합니다. 하지만 테스트 세트에서 이 게이트는 쿼리의 92.2 퍼센트를 9.2 퍼센트 오차로 유지하는데, 이는 예산의 거의 두 배이며, 2 퍼센트 목표는 5.3 퍼센트를 기록합니다. 데이터가 이를 설명해줍니다: Laya는 검증 쿼리의 92.7 퍼센트에서 옳지만 테스트 쿼리에서는 87.8 퍼센트만 옳으므로, 한 샘플에 피팅된 오차 예산은 그와 닮은 트래픽에만 유효합니다. 순위 자체는 건전하여, 가장 확신이 높은 절반의 테스트 답변은 97.8 퍼센트 정확하지만, 오차 목표에는 마진과 실제 트래픽에 대한 주기적 재피팅이 필요합니다. 임계값은 옵션 수 버킷별로 유지됩니다. 하나의 숫자는 두 갈래 질문과 열다섯 갈래 질문 사이에서 전이되지 않기 때문입니다.

Copy Code
@section("9. Out-of-scope traffic: the gate vs. an explicit 'other' option")
def out_of_scope():
    global OOS, OTHER
    OOS = test[test.label == "oos"].sample(150, random_state=0).reset_index(drop=True)
    OTHER = test[~test.label.isin(INTENTS + ["oos"])].sample(150, random_state=0).reset_index(drop=True)
    print(f"  {len(OOS)} out-of-scope queries (e.g. {OOS.text[0]!r})")
    print(f"  {len(OTHER)} in-scope queries from other CLINC domains (e.g. {OTHER.text[0]!r})\n")
 
    groups = {"banking": BTEST.text, "other domains": OTHER.text, "out of scope": OOS.text}
    conf = {g: route(s, INTENTS)[1] for g, s in groups.items()}
    thr = GATE["choice:11+"]
    print(f"  A) 15 intents + the 5% gate (threshold {thr:.3f}): share of each group the gate stops")
    for g in groups:
        print(f"     {g:14s} mean confidence {conf[g].mean():.3f}   abstained {(conf[g] < thr).mean():6.1%}")
 
    with_other = INTENTS + ["not a banking request"]
    print(f"\n  B) 16 options: the 15 intents + 'not a banking request', no gate")
    picks = {g: route(s, with_other)[0] for g, s in groups.items()}
    for g in groups:
        print(f"     {g:14s} routed to 'not a banking request' {(picks[g] == 'not a banking request').mean():6.1%}")
    acc_b = float((picks["banking"] == BTEST.label.values).mean())
    print(f"     banking accuracy with the extra option {acc_b:.3f} (15 intents alone: {OK_FIT.mean():.3f})")
    print("\n  Neither tool is free. The gate needs labelled data and gives up some in-scope coverage;")
    print("  the extra option needs no data but changes the question every intent is scored against.")
    stop = (conf["out of scope"] < thr).mean()
    return f"gate stops {stop:.1%} of out-of-scope queries; 'other' option catches {(picks['out of scope'] == 'not a banking request').mean():.1%}"
 
 
out_of_scope()

프로덕션 트래픽에는 라우터가 설계될 때 고려되지 않은 요청이 포함되므로, 우리는 범위 밖(out-of-scope) CLINC 쿼리 150개와 다른 CLINC 도메인의 쿼리 150개를 추가합니다. 보정된 신뢰도는 이들을 뚜렷이 구분합니다: banking 쿼리는 평균 0.912, 다른 쿼리는 약 0.25이며, 5 퍼센트 게이트는 다른 도메인 쿼리의 89.3 퍼센트와 범위 밖 쿼리의 93.3 퍼센트를 차단하면서 banking 쿼리의 7.8 퍼센트에서는 기권합니다. 대안은 레이블 데이터가 필요 없습니다: 여섯 번째가 아닌 열여섯 번째 옵션, 즉 'not a banking request'가 이들의 80.0 퍼센트와 90.0 퍼센트를 잡아내지만, banking 쿼리의 2.0 퍼센트를 다른 곳으로 라우팅하는 대가를 치르고 banking 정확도를 0.878에서 0.864로 낮춥니다. 새 옵션이 모든 의도가 점수화되는 기준을 바꾸기 때문입니다.

Copy Code
IN_SCOPE = {"in_scope": {"type": "noul", "instructions": "Is this a request about the user's bank account, bills, or payments?"}}
 
 
@section("10. A biased yes/no question, and why temperature cannot fix it")
def biased_noul():
    states = pd.concat([BTEST.text, OTHER.text, OOS.text], ignore_index=True)
    y = np.r_[np.ones(len(BTEST)), np.zeros(len(OTHER) + len(OOS))].astype(bool)
    res = agent.predict_batch(list(states), IN_SCOPE, batch_size=32)
    p = np.array([r["answers"]["in_scope"]["noul"] for r in res])
    from sklearn.metrics import roc_auc_score
    auc = roc_auc_score(y, p)
    print(f"  'Is this a request about the user's bank account, bills or payments?' on 450 banking")
    print(f"  queries and 300 that are not:")
    print(f"    mean P(yes): banking {p[y].mean():.3f}, not banking {p[~y].mean():.3f}   AUROC {auc:.3f}")
    print(f"    at the 0.5 cut: recall {(p[y] >= 0.5).mean():.1%}, specificity {(p[~y] < 0.5).mean():.1%}")
 
    vstates = pd.concat([BVAL.text, val[~val.label.isin(INTENTS + ["oos"])].sample(150, random_state=0).text,
                         val[val.label == "oos"].text], ignore_index=True)
    vy = np.r_[np.ones(len(BVAL)), np.zeros(len(vstates) - len(BVAL))]
    recs = records_from_labeled(agent, [(s, IN_SCOPE, {"in_scope": np.array([1 - t, t])}) for s, t in zip(vstates, vy)])
    t_fit = laya.fit_temperature_map(recs)["temperature"][2]
    print(f"\n  Temperature fitted on {len(recs)} validation answers: T = {t_fit:.2f} (the ceiling is 5.0)")
    print("  Any T leaves the 0.5 cut where it is: dividing two logits by T never changes which is larger.")
    t_ship = agent.temperature_by_options.get("noul:2", agent.temperature[2])
    logits = np.array([r[1] for r in recs])
    p_val = 1 / (1 + np.exp(-(logits[:, 1] - logits[:, 0]) / t_ship))
    cuts = np.linspace(0.01, 0.99, 99)
    bal = [((p_val[vy == 1] >= c).mean() + (p_val[vy == 0] < c).mean()) / 2 for c in cuts]
    cut = float(cuts[int(np.argmax(bal))])
    print(f"\n  What does work is moving the cut. Chosen on validation (best balanced accuracy): {cut:.2f}")
    print(f"    on test at {cut:.2f}: recall {(p[y] >= cut).mean():.1%}, specificity {(p[~y] < cut).mean():.1%}")
    print("\n  The question ranks well and is biased towards 'no'. Temperature scaling repairs a scale,")
    print("  not an offset. Treat P(yes) as a score, and pick its cut on labeled data like any other.")
    globals()["IN_SCOPE_CUT"] = cut
    return f"recall at 0.5 {(p[y] >= 0.5).mean():.1%} -> {(p[y] >= cut).mean():.1%} at a validated cut of {cut:.2f}"
 
 
biased_noul()

전용 yes/no 질문이 자연스러운 범위 내(in-scope) 확인처럼 보이므로, 우리는 각 메시지가 사용자의 은행 계좌, 청구서 또는 결제에 관한 것인지 묻습니다. AUROC 0.945로 순위는 좋지만 no 쪽으로 편향되어 있습니다: banking 쿼리는 평균 확률이 0.361에 불과하여, 기본 0.5 컷에서는 이들 중 28.9 퍼센트만 인식합니다. 550개의 검증 답변에 temperature를 피팅하면 상한 5.0에 도달하며 그 컷에서 아무것도 바뀌지 않는데, 두 로짓을 어떤 temperature로 나누더라도 어느 쪽이 더 큰지는 변하지 않기 때문입니다. Temperature 스케일링은 스케일을 수리할 뿐 오프셋은 수리하지 못합니다. 실제로 효과가 있는 것은 확률을 점수로 취급하고 레이블이 달린 데이터에서 그 컷을 선택하는 것입니다: 검증 분할에서 선택한 0.09의 컷은 테스트 세트에서 92.0 퍼센트 재현율과 83.0 퍼센트 특이도를 제공합니다.

Copy Code
from typing import Literal, Optional
from pydantic import BaseModel, Field
 
 
class BankingRequest(BaseModel):
    intent: Optional[Literal[tuple(INTENTS)]] = Field(description=ASK)
    in_scope: bool = Field(description=IN_SCOPE["in_scope"]["instructions"])
 
 
@section("11. Typed decisions from a pydantic schema")
def schema_decisions():
    planned = laya.structured.questions_from_pydantic(BankingRequest)
    crit = planned["intent"]["criteria"]
    print(f"  questions_from_pydantic(BankingRequest): 'intent' -> {planned['intent']['type']} with "
          f"{len(crit)} options, criteria values {set(crit.values())};  'in_scope' -> {planned['in_scope']['type']}")
    print("  A Literal becomes a choice over bare names (what step 5 found works best here); a bool becomes")
    print("  a yes/no question that the projection cuts at 0.5.\n")
    demo = pd.concat([BTEST.iloc[[0, 120, 300]], OOS.iloc[[0, 1]]], ignore_index=True)
    out = laya.decide_batch(agent, list(demo.text), schema=BankingRequest, return_details=True,
                            min_confidence=GATE, batch_size=8)
    for (text, label), d in zip(zip(demo.text, demo.label), out):
        typed = BankingRequest(**d.values)
        p_yes = d.probabilities["in_scope"]["true"]
        print(f"  {text[:46]!r:50s} truth {label}")
        print(f"      -> {typed!r}   P(in_scope) {p_yes:.2f}, at our cut {p_yes >= IN_SCOPE_CUT}")
    print("\n  min_confidence=GATE turns a gated answer into None, so 'intent=None' means 'ask a human',")
    print("  and the pydantic model still validates. in_scope is decided at a fixed 0.5 inside decide();")
    print("  read d.probabilities['in_scope']['true'] and apply the cut from step 10 instead.")
    nones = sum(BankingRequest(**d.values).intent is None for d in out)
    return f"{len(out)} typed objects, {nones} intents gated to None"
 
 
schema_decisions()

마지막으로, 우리는 이것을 pydantic 스키마를 통해 애플리케이션 코드에 연결합니다. laya.decide_batch는 Literal 필드를 순수 이름(bare names)에 대한 choice 질문으로(단계 5에서 보았듯이 여기서는 더 나은 표현입니다), bool을 yes/no 질문으로 변환한 뒤, 답변을 검증된 모델 인스턴스로 다시 투영합니다. 단계 8의 게이트를 min_confidence로 전달하면 불확실한 의도가 None으로 바뀌며, Optional 필드가 이를 받아들이므로 None은 명시적인 사람에게 물어보기(ask-a-human) 신호가 됩니다. 두 개의 범위 밖 쿼리는 intent=None으로 돌아옵니다. 하지만 불리언은 decide 내부에서 고정된 0.5로 잘리므로, 사기 보고는 확률 0.29에서 in_scope=False로 표시됩니다. 결과 세부 정보에서 확률을 읽고 단계 10의 컷을 적용하면 올바른 답을 얻습니다.

코드 복사
banner("SUMMARY")
for name, res in RESULTS.items():
    print(f"  {name:<76s}  {res}")
print("""
What to carry over
 - Pin the checkpoint (laya.PINNED_REVISIONS) and know its shipped temperatures before trusting a
   probability: 15-option questions were sharpened by a clamped 0.5.
 - Measure criteria wording and option order on labeled data. Bare names beat our descriptions.
 - Fit temperatures on held-out data, and install only the bucket you measured: fit_temperatures()
   replaces the whole map, including question types you did not fit.
 - Gate per option-count bucket with fit_abstention_thresholds, and leave margin below the target.
 - A yes/no question can rank well and still be biased; choose its cut on labeled data.
Where to go next
 - Fine-tuning: laya.train on the repository's main branch (not yet in the 0.3.27 wheel) and the
   repo's Kaggle / Apple-silicon notebooks train with RLCD and refit temperatures.
 - Other languages: laya.Router() detects the script and routes to laya-multilingual, which ships with
   no fitted temperatures at all.
 - Serving: pip install "laya[serve]" for an HTTP server; laya[mcp] for an MCP tool server.
 - Docs: nandhakishorm.github.io/laya    Code: github.com/NandhaKishorM/laya
""")

요약은 각 단계의 한 줄 결과와 앞으로도 유지할 만한 습관들을 출력합니다: 체크포인트를 고정하고 함께 배포된 temperature를 읽어오기, 라벨링된 데이터에서 기준 문구와 옵션 순서 테스트하기, 홀드아웃 데이터에서 temperature를 피팅하고 측정한 버킷만 설치하기, 마진을 두고 옵션 수별 버킷으로 게이트하기, 라벨링된 데이터에서 예/아니오 컷을 선택하기. 마지막으로 다음 단계를 안내합니다: 파인튜닝은 저장소의 main 브랜치에 있는 laya.train에 존재하지만 아직 0.3.27 휠에는 포함되지 않았으며, laya.Router 뒤에 있는 다국어 체크포인트, 그리고 서빙입니다.

결론적으로, Laya는 약속한 것의 대부분을 전달합니다: 생성 토큰 없이 단 한 번의 순전파로 여러 개의 타입화된 질문에 답하고, 학습 데이터 없이 베어 이름 라우터가 열다섯 개의 실제 은행 인텐트에서 0.878을 달성했는데 이는 TF-IDF 분류기가 맞추려면 인텐트당 열 개에서 서른 개 사이의 라벨링된 예제가 필요한 수준이며, 보정된 신뢰도가 범위 내 트래픽과 범위 외 트래픽을 충분히 구분하여 범위 외 질의 열 개 중 아홉 개 이상을 차단했습니다. 하지만 그 확률이 실제로 얼마나 가치 있는지는 라이브러리가 사용자에게 맡기는 작업에 달려 있으며, 여러 기본 설정은 잘못된 방향을 가리킵니다. 옵션이 열한 개 이상인 질문에 대한 기본 제공 temperature는 완화시키지 않고 오히려 날카롭게 만들고, 한 줄짜리 보정 호출은 한 번도 본 적 없는 질문 유형의 temperature를 지워버리며, 검증 데이터에 피팅된 오류 예산은 그 데이터와 비슷해 보이는 트래픽에만 적용되고, 예/아니오 질문은 순위는 잘 나오지만 0.5의 잘못된 쪽에 놓일 수 있는데 이는 어떤 temperature로도 해결할 수 없고 스키마 투영이 이를 하드코딩합니다. 이 각각의 문제에는 위에서 보여준 몇 줄짜리 수정법이 있으며, 그렇지 않으면 조용히 통과될 수 있습니다. 실용적인 교훈은 이 라이브러리의 모델 카드가 직접 말하고 있지만 기본 설정 때문에 잊기 쉬운 바로 그것입니다: 보정된 의사결정 모델이란 여러분이 여러분의 라벨로, 여러분의 질문에 대해 직접 보정한 모델입니다.


전체 코드는 여기에서 확인하세요. 이 프로젝트의 연구자에게 모든 크레딧이 돌아갑니다. 또한 언제든지 저희를 Twitter 에서 팔로우하고 150k+ML SubReddit 에 가입하고 our Newsletter을 구독하는 것도 잊지 마세요. 잠깐! 텔레그램을 사용하고 계신가요? 이제 텔레그램에서도 저희와 함께하실 수 있습니다.

[Sponsored] 웹은 대부분의 에이전트가 빠뜨린 유일한 API입니다. 데이터베이스, 캘린더, 저장소에는 API가 있습니다. 열린 웹에는 대부분 없습니다. TinyFish MCP 서버 는 모든 MCP 클라이언트에 네 가지 도구를 제공합니다: TinySearch, TinyFetch(JavaScript를 포함한 전체 페이지를 마크다운으로), 로그인과 폼을 위한 TinyBrowser, 그리고 다단계 작업을 위한 TinyAgent. Search와 Fetch는 무료입니다.

이 글 A Developer’s Guide to Laya: Zero-Shot Decisions and Calibration 은 다음에서 처음 게시되었습니다 MarkTechPost.

원문 출처

MarkTechPost

내용 안내

원문 발행 및 권리는 출처에 있습니다.

기계 번역 · 원문을 참고하세요