MarkTechPost

Google Research RRSIガイド: 自己改善型AIエージェントをマスターする

Google ResearchのRRSI(Regularized Recursive Self-Improvement)に関する包括的なコーディングガイドを通じて、ノイズバンド、コストルール、リークスクリーニングがどのように安全で効率的な自己改善型AIエージェントを実現するかを探ります。本記事「Google Research RRSI Guide: Mastering Self-Improving AI…

このチュートリアルでは、次を実装します。 RRSI(Regularized Recursive Self-Improvement)。これは、凍結されたモデルを前提に、LLMエージェントが自身のハーネス、プロンプト、ツール、メモリ、制御フロー、サブエージェントを書き換えられるようにする手法であり、ハーネスが進化対象のタスクに過適合しないようにします。完全なRRSIループはVertex AI上のClaude Opusで編集案を作成し、Dockerベンチマーク内でスコアを評価します。これは無料のノートブックでは実行できません。しかし、RRSIのうち論文のアイデアを実際に担う部分、すなわちどの編集案を採用するかを決めるルールは、プレーンなPythonで書かれており、われわれはこれを直接動かします。公式リポジトリからパッケージをインストールし、そのエスティメーター、較正されたノイズバンド、選択アルゴリズムの両分岐、アニールされる編集予算、決定論的なリークスクリーニング、編集履歴を順に見ていき、その後、シミュレートされたエージェントをRRSI独自のDomainインターフェースに接続します。シミュレート環境は自分たちで構築したものなので、すべての編集の真の効果が分かり、RRSIの判断を正解データと突き合わせて監査し、単に最もスコアの高いものを保持する非正規化探索と比較できます。

コードをコピー
import os
import sys
import json
import math
import copy
import random
import tempfile
import textwrap
import traceback
import subprocess
import statistics as st
from pathlib import Path
 
RESULTS = {}
 
 
def banner(title):
    print("\n" + "=" * 78)
    print(title)
    print("=" * 78)
 
 
def section(name):
    def wrap(fn):
        def run(*a, **kw):
            banner(name)
            try:
                out = fn(*a, **kw)
                RESULTS[name] = out if isinstance(out, str) else "ok"
                return out
            except Exception as e:
                RESULTS[name] = f"SKIPPED / FAILED -> {type(e).__name__}: {e}"
                print(f"\n[!] {name} did not complete: {type(e).__name__}: {e}")
                traceback.print_exc(limit=3)
                return None
        return run
    return wrap
 
 
banner("1. Install RRSI and map the paper onto the code")
subprocess.run([sys.executable, "-m", "pip", "install", "-q", "git+https://github.com/google-research/rrsi.git@be50316e1db05914068a973f322770ef08ed7ba1"], check=True)
 
from importlib.metadata import version
from rrsi.config import RRSIConfig
from rrsi.evaluate import TaskResult, EvalResult, aggregate, evaluate
from rrsi.calibrate import calibrate
from rrsi.selection import Candidate, cost_rule, judge, select_round
from rrsi.schedule import edit_budget, budget_table
from rrsi.history import History, stall_flag, exploration
from rrsi.components import K, K_STR, normalize, novelty
from rrsi.critic import precheck, review
from rrsi.domain import Domain
 
print(f"  rrsi {version('rrsi')}  |  anthropic {version('anthropic')}  |  Python {sys.version.split()[0]}")
print("\n  RRSI evolves an agent's HARNESS (prompts, tools, memory, control flow, sub-agents) around a")
print("  frozen model. The full loop drafts edits with Claude Opus on Vertex AI and scores them in Docker")
print("  benchmarks. The part that decides which edits to KEEP is plain Python, and that is what we drive:")
for symbol, where in [
    ("S_hat, C_hat       Eq. (estimate)", "rrsi.evaluate.aggregate"),
    ("delta              noise band", "rrsi.calibrate.calibrate"),
    ("Algorithm 2        floor + cost rule", "rrsi.selection.judge / select_round"),
    ("b_t                Eq. (anneal)", "rrsi.schedule.edit_budget"),
    ("Critic             leakage screen", "rrsi.critic.precheck / review"),
    ("L_t, g_t, B_t      history, yield, prune", "rrsi.history.History"),
    ("sigma_t, U_t       stall + exploration", "rrsi.history.stall_flag / exploration"),
    ("nu                 structural novelty", "rrsi.components.novelty"),
]:
    print(f"    {symbol:40s} -> {where}")
 
CFG = RRSIConfig()
print(f"\n  paper defaults: T={CFG.T} rounds, k={CFG.k} trials/task, m={CFG.m} candidates/round,"
      f" b in [{CFG.b_min},{CFG.b_max}]")
print(f"                  beta0={CFG.beta0}  beta1={CFG.beta1}  w_s={CFG.w_s}  w_c={CFG.w_c}  w_n={CFG.w_n}"
      f"  delta_z={CFG.delta_z}")
print("\n  Nothing below needs an API key, a GPU or a dataset download.")

RRSIはgoogle-researchリポジトリからインストールします。このパッケージはPyPIにないため、このノートブックが執筆された時点のコミットに固定します。唯一の依存関係はAnthropicクライアントで、これは探索ロールがClaudeを呼び出すために使うものですが、われわれは一度も使いません。次に、リポジトリ自体が文書化している、論文の記号とそれを実装する関数との対応を出力します。evaluateにある経験的スコアとコスト推定、calibrateにあるノイズバンド、selectionにあるAlgorithm 2、scheduleにあるアニールされる編集予算、criticにあるリークスクリーニング、そしてhistoryにあるyield、prune、stall、探索のサマリーを伴う編集履歴です。RRSIConfigは論文のハイパーパラメータを保持し、以下のすべての関数は実際のループとまったく同じようにこれを受け取ります。

コードをコピー
@section("2. Evaluate(H): a score and a cost, and why a crash counts as zero")
def estimator():
    base = {
        "task_000": TaskResult(rewards=[1, 1], tokens=[11_800, 12_400]),
        "task_001": TaskResult(rewards=[1, 0], tokens=[15_100, 14_600]),
        "task_002": TaskResult(rewards=[0, 0], tokens=[21_000, 19_500]),
    }
    ev = aggregate("H0", 2, base)
    print(f"  three tasks x k=2 trials -> S_hat = {ev.S:.3f}   C_hat = {ev.C:,.0f} tokens/trial"
          f"   ({ev.n_expected} trials expected, {ev.missing} missing)")
 
    crashy = dict(base)
    crashy["task_002"] = TaskResult(rewards=[0.0, 0.0], tokens=[None, None], missing=2)
    ev_crash = aggregate("crashy", 2, crashy)
    dropped = {t: r for t, r in crashy.items() if not r.missing}
    naive = sum(sum(r.rewards) for r in dropped.values()) / sum(len(r.rewards) for r in dropped.values())
    print("\n  A candidate crashes on the hardest task instead of failing it:")
    print(f"    an estimator that drops missing trials reports  {naive:.3f}   <- looks like a gain")
    print(f"    RRSI's aggregate (missing = 0, full denominator) {ev_crash.S:.3f}   <- no reward for crashing")
    print(f"    ...and C_hat uses only recorded token counts: {ev_crash.C:,.0f}")
 
    rubric = {"memo": TaskResult(rewards=[0.5, 1.0], weights=[10, 10]),
              "brief": TaskResult(rewards=[0.0, 0.0], weights=[90, 90])}
    ev_w = aggregate("rubric", 2, rubric)
    print("\n  Weighted rewards (Harvey LAB style: weight = number of rubric criteria):")
    print(f"    mean of per-task means = {st.mean(r.mean for r in rubric.values()):.3f}"
          f"   vs   RRSI's S_hat = {ev_w.S:.3f} (fraction of all criteria passed)")
    return f"crash scored {ev_crash.S:.3f} under RRSI vs {naive:.3f} if dropped"
 
 
estimator()

RRSIはハーネスごとに2つの数値を測定します。すべてのタスクのすべてのトライアルにわたる平均報酬であるSと、トライアルあたりの平均ポリシートークン数であるCです。TaskResultは1タスクのトライアルを記録し、それらを集計します。どのエージェント評価にも応用できる重要な点は、欠損トライアルの扱われ方です。候補が最も難しいタスクでクラッシュしたとき、欠損トライアルを除外するエスティメーターは0.750を報告し、クラッシュが改善のように見えてしまいます。一方、RRSIは各欠損トライアルを分母をそのままにゼロ報酬として数え、以前と同じ0.500を報告するため、候補は自分が難しいと感じたトライアルを破壊することで良く見えることはできません。重み付き報酬は、Harvey LABのようなルーブリック評価型スイートに対応し、この場合Sはタスクごとの平均の平均ではなく、すべての基準を通過した割合になります。

コードをコピー
FAMILIES = ["parse", "search", "edit", "test"]
H0 = {"skill": {f: 0.0 for f in FAMILIES}, "memo": [], "cost": 1.0}
 
 
def make_world(seed, n_evolve, n_heldout=80):
    """Tasks with a family and a difficulty. Evolve ids are task_NNN, held-out ids held_NNN."""
    r = random.Random(seed)
    world = {f"task_{i:03d}": (FAMILIES[i % 4], r.gauss(0, 1)) for i in range(n_evolve)}
    world.update({f"held_{i:03d}": (FAMILIES[i % 4], r.gauss(0, 1)) for i in range(n_heldout)})
    return world
 
 
def p_success(h, world, task):
    """The frozen policy: a logistic in harness skill minus task difficulty, or 0.97 if memorised."""
    if task in h["memo"]:
        return 0.97
    family, difficulty = world[task]
    return 1 / (1 + math.exp(-(0.3 + h["skill"][family] - difficulty)))
 
 
def run_trials(h, world, ids, k, rng):
    return {t: TaskResult(rewards=[float(rng.random() < p_success(h, world, t)) for _ in range(k)],
                          tokens=[round(12_000 * h["cost"] * math.exp(rng.gauss(0, 0.08))) for _ in range(k)])
            for t in ids}
 
 
def true_score(h, world, ids):
    return sum(p_success(h, world, t) for t in ids) / len(ids)
 
 
def calibrated_delta(world, k, seed, repeats=3):
    """delta from `repeats` independent evaluations of the SAME harness H0."""
    ids = [t for t in world if t.startswith("task_")]
    rng = random.Random(10_000 + seed)
    evals = [aggregate(f"base{j}", k, run_trials(H0, world, ids, k, rng)) for j in range(repeats)]
    return calibrate(evals, z=CFG.delta_z, reps=200), evals
 
 
@section("3. The noise band: how far one harness's score moves between two evaluations")
def noise_band():
    world = make_world(0, 40)
    ids = [t for t in world if t.startswith("task_")]
    rng = random.Random(1)
    scores = [aggregate(f"H0#{j}", 2, run_trials(H0, world, ids, 2, rng)).S for j in range(6)]
    print(f"  H0 evaluated 6 times on 40 tasks x k=2: {[round(s, 3) for s in scores]}")
    print(f"  true expected score {true_score(H0, world, ids):.3f}; spread of the estimates"
          f" {max(scores) - min(scores):.3f}. Nothing about the harness changed.")
 
    print(f"\n  calibrate() turns repeated evaluations into delta = z * sd(null dS), z = {CFG.delta_z}:")
    print(f"  {'evaluator':>12s} {'trials':>7s} {'delta':>8s}   method")
    deltas = {}
    for n, k in [(40, 2), (100, 4), (400, 8)]:
        cal, _ = calibrated_delta(make_world(0, n), k, seed=0)
        deltas[f"{n}x{k}"] = cal["delta"]
        print(f"  {f'{n} x k={k}':>12s} {n * k:>7d} {cal['delta']:>8.4f}   {cal['method']}")
    print("\n  The paper's calibrated bands are 0.017 (coding), 0.004 (workspace) and 0.020 (engineering).")
    print("  A gain smaller than delta is indistinguishable from re-running the same harness, and")
    print("  Algorithm 2 treats it that way. Remember the first row: it becomes the lesson of step 11.")
    return "delta " + ", ".join(f"{k}={v:.3f}" for k, v in deltas.items())
 
 
noise_band()

どんなルールも本当の改善と運を切り分ける前に、1つのハーネスのスコアがそれ自身でどれほど動くかを知る必要があります。われわれは小さなシミュレートエージェントを構築します。その各タスクでの成功率は、ハーネスのスキルからタスクの難易度を引いた値のロジスティック関数であり、変更していない初期ハーネスを、40タスク×2トライアルで6回評価すると、何も変わっていないのにスコアは0.113ぶんばらつきました。calibrateは同一ハーネスの繰り返し評価をdelta、すなわち2回の実行間の差の標準偏差の2倍に変換します。80トライアルではdeltaは約0.108、3,200トライアルでは約0.013まで下がります。これは論文がそのインスタンスについて報告している範囲(0.004から0.020)です。選択の目的上、deltaより小さい改善は、同じハーネスの再実行と区別がつきません。

コードをコピー
def ev_at(S, C, job, n=200):
    """An EvalResult with exactly score S (in steps of 1/n) and C tokens per trial."""
    hits = round(S * n)
    return aggregate(job, 1, {f"task_{i:03d}": TaskResult(rewards=[1.0 if i < hits else 0.0], tokens=[C])
                              for i in range(n)})
 
 
INCUMBENT = ev_at(0.630, 10_000, "incumbent")
S_STAR, DELTA = 0.640, 0.020
 
 
@section("4. Algorithm 2, branch by branch: the floor, the cost rule and the noise band")
def algorithm_2():
    print(f"  incumbent S = {INCUMBENT.S:.3f}, C = {INCUMBENT.C:,.0f}   best ever S* = {S_STAR}   delta = {DELTA}")
    print(f"  floor = S* - delta = {S_STAR - DELTA:.3f}\n")
    cases = [
        ("A  slipped below the floor", 0.600, 10_000, ["prompt"], None),
        ("B  real gain, pays for its tokens", 0.690, 12_000, ["prompt"], None),
        ("C  real gain, far too expensive", 0.660, 25_000, ["subagent"], None),
        ("D  slightly WORSE but cheaper", 0.625, 8_000, ["context_mgmt"], None),
        ("E  in the band but costlier", 0.640, 11_500, ["prompt"], None),
        ("F  in the band, new sub-agent", 0.630, 10_000, ["subagent"], None),
        ("F' in the band, prompt tweak", 0.630, 10_000, ["prompt"], None),
        ("G  like B, breaks a domain guard", 0.690, 12_000, ["prompt"], ["valid-output rate fell"]),
    ]
    rows = []
    for label, S, C, comps, guards in cases:
        cand = Candidate(label[:2].strip(), [{"id": "C1", "component": c} for c in comps], ev=ev_at(S, C, label))
        dec = judge(cand, INCUMBENT, S_STAR, DELTA, CFG, incumbent_counts={}, guards=guards)
        rows.append((label, dec))
        print(f"  {label:36s} S={S:.3f} C={C:>6,}  {'ADMIT ' if dec.admissible else 'reject'}")
        print(textwrap.indent(textwrap.fill(dec.reason, 88), " " * 6))
 
    print("\n  Three rules, in the order RRSI applies them:")
    print("    1. never fall below the best score ever seen, minus the noise band        (A)")
    print("    2. a gain bigger than delta must pay for any extra tokens: dC <= 0.10 + 40*dS  (B, C)")
    print("    3. inside the band scores are a tie, so prefer the cheaper harness, and let a")
    print("       never-tried structural component break the tie                          (D, E, F, F')")
    print("  D is the surprising one: RRSI admits a harness that scored LOWER than the incumbent.")
    return f"{sum(d.admissible for _, d in rows)}/{len(rows)} admissible; D admitted at dS={rows[3][1].delta_S:+.3f}"
 
 
algorithm_2()

Algorithm 2は純粋な関数として実装されているため、候補を与えてその判断理由をそのまま読み取ることができます。現行をS 0.630、トークン数10,000、史上最高スコア0.640、deltaを0.020に固定し、8つの候補をjudgeに渡します。フロア、すなわちそれまでに見た最高スコアからdeltaを引いた値を下回る候補は即座に拒否されます。deltaを超える改善は、相対コスト変化が0.10に改善幅の40倍を足した値を下回るというルールの下で、追加トークンのコストに見合わなければなりません。トークン+20%での+6ポイントの改善は認められ、トークン+150%での+3ポイントの改善は認められません。バンド内ではスコアは同点として扱われ、改善幅の100倍からコスト変化の15倍を引いた整形スコアに、未採用の構成要素に対する小さなボーナスを加えたものが決め手になります。これにより、RRSIは現行よりスコアが低いもののトークンコストが20%少ない候補Dを採用し、また新しいサブエージェントは、同等のスコアのプロンプト調整では解消されない同点を打破します。ドメインガードはスコアにかかわらず拒否権を行使します。

コードをコピー
@section("5. One round of selection: the highest score does not always win")
def one_round():
    c_hi = Candidate("C", [{"id": "C1", "component": "subagent"}], ev=ev_at(0.660, 25_000, "C"))
    d_lo = Candidate("D", [{"id": "C1", "component": "context_mgmt"}], ev=ev_at(0.625, 8_000, "D"))
    leak = Candidate("X", [{"id": "C1", "component": "memory"}], gate_failure="critic_reject")
    winner, decisions = select_round([c_hi, d_lo, leak], INCUMBENT, S_STAR, DELTA, CFG, incumbent_counts={})
    for d in decisions:
        print(f"  {d.variant}: {'admissible' if d.admissible else 'rejected  '}  {d.reason[:92]}")
    new_star = max(S_STAR, winner.ev.S)
    print(f"\n  winner: {winner.variant}  (S {winner.ev.S:.3f}, C {winner.ev.C:,.0f})")
    print(f"  the new incumbent scores {winner.ev.S - INCUMBENT.S:+.3f} vs the old one and costs"
          f" {(winner.ev.C - INCUMBENT.C) / INCUMBENT.C:+.0%} tokens; S* stays {new_star:.3f}")
    print("\n  C scored highest and still lost: its +3pp does not pay for +150% tokens. D moved the")
    print("  incumbent DOWN inside the noise band because it is 20% cheaper. Because S* only ever")
    print("  rises, the floor never follows the incumbent down, so a chain of 'cheaper but slightly")
    print("  worse' swaps cannot walk the score away. X never reached evaluation at all.")
    return f"winner {winner.variant} at dS={winner.ev.S - INCUMBENT.S:+.3f}, dC={(winner.ev.C - INCUMBENT.C) / INCUMBENT.C:+.0%}"
 
 
one_round()

select_round はラウンド内のすべての候補に judge を適用し、最もスコアの高い許容可能な候補を保持します。ここに、高コストな候補、わずかに劣る代わりに低コストな候補、そして critic がすでに拒否した候補を与えます。最高スコアの候補は、3ポイントの改善では150%増のトークンコストに見合わないために失われ、critic に拒否された候補は評価にすら到達せず、勝者は現行ハーネスのスコアを0.5ポイント下げる一方でトークンコストを5分の1に削減します。これを安全に保つ詳細が S* であり、それは単調に上昇するだけです。下限は現行ハーネスではなく、これまでに測定された最高スコアに固定されているため、より安いがわずかに劣る候補への交換の連鎖が、多くのラウンドにわたってスコアを少しずつ押し下げていくことはできません。

コードをコピー
@section("6. The proposal side: an edit budget that anneals from 4 toward 1")
def edit_budget_schedule():
    table = budget_table(CFG.T, CFG.b_min, CFG.b_max)
    print("  b_t = ceil(b_min + (b_max - b_min) * (1 + cos(pi t / T)) / 2)")
    print(f"  T={CFG.T}, b in [{CFG.b_min}, {CFG.b_max}]:")
    print("    t   : " + " ".join(f"{t:>2d}" for t in range(CFG.T)))
    print("    b_t : " + " ".join(f"{b:>2d}" for b in table))
    print(f"    round {CFG.T} (after the run) -> {edit_budget(CFG.T, CFG.T, CFG.b_min, CFG.b_max)}")
    print("\n  Early candidates may bundle up to 4 coordinated edits. Note the ceil(): the cosine term is")
    print(f"  only exactly zero at t = T, so inside a {CFG.T}-round run the budget bottoms out at"
          f" {min(table)}, not {CFG.b_min}.")
    print("  Every edit in a bundle inherits the bundle's single measurement, so the shrinking budget is")
    print("  what makes late history attributable to fewer components. It caps how MANY edits ride")
    print("  together, never WHICH mechanisms the harness may eventually contain.")
    return "budget " + "".join(str(b) for b in table)
 
 
edit_budget_schedule()

提案側は、どの編集が保持されるかではなく、編集がどのように起草されるかを正規化します。edit_budget は論文のアニールされた L0 予算を実装しています。これは b_max から b_min へのコサインスケジュールで、デフォルトでは最初の8ラウンドは候補あたり最大4つの編集、次の5ラウンドは3つ、最後の7ラウンドは2つまでの協調的な編集を許可します。式の中の天井関数により、予算が最小値の1に到達するのは t = T のとき、すなわち実行終了の1ステップ後であり、方程式を読む際に見落としやすい点です。バンドル内のすべての編集はバンドルの1回の測定を引き継ぐため、予算の縮小こそが、実行後半の履歴をより少ない構成要素に帰属させる仕組みになっています。予算は一緒に移動できる編集の数を制限するだけであり、ハーネスが最終的にどのメカニズムを含みうるかを制限することはありません。

コードをコピー
CRITIC_PATTERNS = [
    (r"\btask_\d{3}\b", "hard-codes an evolve-set task id"),
    (r"expected_output|grader|rubric\[", "reads the grader or the expected answer"),
]
COMPONENT_SIGNALS = [
    ("control_flow", [r"\bretry\(", r"max_attempts"]),
    ("context_mgmt", [r"compress_context", r"keep_last"]),
    ("config", [r"CONFIG\["]),
]
 
 
class ScreenOnly(Domain):
    name = "screen"
    critic_patterns = CRITIC_PATTERNS
    briefs = {"critic": "A coding agent harness."}
 
 
@section("7. The critic's deterministic layer, and how edits are tagged")
def critic_and_tags():
    dom = ScreenOnly()
    diffs = {
        "memorise answers": "+ memory = Memory('answers')\n+ memory.remember('task_007', cached_patch)",
        "peek at the grader": "+ if os.path.exists('/grader/expected_output.txt'): return read_it()",
        "leaked credential": "+ api_key = 'sk-live-0123456789abcdefghijkl'",
        "empty diff": "   ",
        "general retry rule": "+ for attempt in range(max_attempts): result = retry(step)",
    }
    for label, diff in diffs.items():
        try:
            verdict = review(dom, diff, summary=label, targets_mode="evolve")
            print(f"  {label:20s} -> {verdict['verdict']:6s} {verdict['reasons']}")
        except ZeroDivisionError as e:
            print(f"  {label:20s} -> passed the deterministic layer; the LLM layer raised"
                  f" ZeroDivisionError: {e}")
    print("\n  Gotcha: with RRSI_VERTEX_PROJECTS unset, rrsi.llm.generate computes `x % len(projects)`")
    print("  outside its try block, so the helpful 'set RRSI_VERTEX_PROJECTS' error is never reached.")
    print("  A clean diff is supposed to go to Claude for an intent review; here there is no Claude.")
 
    print("\n  Every edit is tagged with the component it touches, and a tag needs evidence in the diff:")
    tag_cases = [
        ("skill", "+ \"Remember to run the tests before finishing.\""),
        ("control_flow", "+ for attempt in range(max_attempts): result = retry(step)"),
        (None, "+ review = subcall('reviewer', transcript)"),
        ("memory", "+ context = compress_context(context, keep_last=8)"),
    ]
    for declared, diff in tag_cases:
        tag = normalize(declared, diff, COMPONENT_SIGNALS)
        print(f"    declared {str(declared):13s} -> tagged {tag:13s} {diff[2:60]!r}")
    print("  A proposer cannot label a prompt tweak as a new 'skill' to look novel: without evidence")
    print("  the tag falls back to what the diff actually is.")
    counts = {"prompt": 3, "subagent": 1}
    print(f"\n  novelty(nu) counts STRUCTURAL components {K_STR} the incumbent has never accepted.")
    print(f"  against an incumbent with accepted edits {counts}:")
    for comps in (["memory"], ["subagent"], ["prompt", "client_tool", "memory"]):
        print(f"    {str(comps):38s} nu = {novelty(comps, counts)}")
    return "precheck rejected 4/5 diffs without an LLM call"
 
 
critic_and_tags()

critic は評価コストが一切費やされる前に、2層構造ですべての候補 diff をふるいにかけます。第1層は、汎用的なクレデンシャルパターンとドメイン固有の denylist に対する決定論的な事前チェックです。evolve-set のタスクIDと grader パスのパターンにより、task_007 の答えを暗記する diff、期待される出力を読み取る diff、APIキーを漏洩する diff、空の diff を、モデルを一切呼び出さずに拒否します。クリーンな diff は第2層である Claude による意図レビューに進みますが、ここでノートブックは実際の落とし穴を明らかにします。RRSI_VERTEX_PROJECTS が未設定のとき、rrsi.llm.generate は try ブロックの外側でプロジェクト数を法としたインデックスを計算し ZeroDivisionError を送出するため、コード内の役に立つ設定エラーメッセージには決して到達しません。また、編集へのタグ付けの仕方も見ていきます。normalize は、diff に証拠がある場合にのみ宣言されたコンポーネントを保持するため、プロンプトの微調整が新規性ボーナスを得るために新しいスキルを装うことはできず、誤ってラベル付けされたコンテキスト管理の変更は、実際の内容どおりにタグ付けされます。

コードをコピー
@section("8. The edit history: what was tried, what paid off, what to prune")
def edit_history():
    with tempfile.TemporaryDirectory() as tmp:
        h = History(Path(tmp) / "history.jsonl")
        log = [
            (0, "A", [("prompt", "tell the agent to read the failing test first")], "ACCEPTED", 0.030, 0.02, True, 0.66),
            (1, "A", [("prompt", "ask for a plan before editing")], "REJECTED", -0.010, 0.05, False, 0.65),
            (1, "B", [("subagent", "add a reviewer sub-agent"), ("memory", "persist lint rules")],
             "REJECTED", -0.020, 0.40, False, 0.64),
            (2, "A", [("config", "raise the step limit")], "ACCEPTED", 0.005, -0.03, True, 0.665),
            (3, "A", [("prompt", "shorter system prompt")], "LOST", -0.001, -0.10, False, 0.664),
            (4, "B", [("memory", "cache task_014 solution")], "critic_reject", None, None, False, None),
            (5, "A", [("prompt", "stricter output format")], "REJECTED", -0.004, 0.00, False, 0.661),
        ]
        for t, v, edits, outcome, dS, dC, acc, S in log:
            h.append_candidate(t, v, [{"id": f"C{i + 1}", "component": c, "hypothesis": hyp}
                                      for i, (c, hyp) in enumerate(edits)],
                               outcome, dS, dC, acc, S, 12_000 if S else None, diff=None)
        t_now = 6
        print(f"  {len(h.records())} per-edit records from {len(log)} candidates"
              f" (the two-edit bundle in round 1 wrote two records with ONE measurement)")
        print(f"  T_t, tried components : {sorted(h.tried())}   (the critic-rejected edit is not 'tried')")
        g_t = h.yield_g(t_now, CFG.n_prune)
        print("  g_t, best gain in the last n_prune rounds: "
              + ", ".join(f"{c} {'none measured' if g == -math.inf else f'{g:+.3f}'}" for c, g in sorted(g_t.items())))
        prune = h.prune_set(t_now, CFG.n_prune)
        print(f"  B_t, prune set        : {[p['component'] for p in prune]}")
        for p in prune:
            if p["accepted_edits_in_incumbent"]:
                print(f"      {p['component']}: still in the incumbent but no recent gain ->"
                      f" {[e['hypothesis'] for e in p['accepted_edits_in_incumbent']]}")
 
        trajectory = [0.630, 0.660, 0.660, 0.665, 0.665, 0.665, 0.665]
        sigma = stall_flag(trajectory, t_now, CFG.w, DELTA)
        ex = exploration(t_now, sigma, h.tried(), CFG.m_draft)
        print(f"\n  S over rounds {trajectory}: moved {trajectory[t_now] - trajectory[t_now - CFG.w]:+.3f}"
              f" in the last w={CFG.w} rounds -> sigma_t = {sigma}")
        print("  what the proposer is told next round:")
        print(textwrap.indent(textwrap.fill(ex["text"], 84), "    "))
    print("\n  The proposer sees this history, so a falsified hypothesis ('ask for a plan first', -1pp)")
    print("  is not redrawn, and a stalled run is pushed toward components it has never touched.")
    return f"prune set {[p['component'] for p in prune]}, stall flag {sigma}"
 
 
edit_history()

履歴は編集ごとに1件の JSONL レコードを書き込み、ループはそこから4つの要約を導出します。tried セットは、測定されたことがないため、critic が破棄した編集を除外します。直近の利得 g_t は直近の n_prune ラウンド内のコンポーネントあたりの最高の測定利得であり、直近の利得が正でないすべてのコンポーネントは、現行ハーネスにまだ残っているそのコンポーネントの仕組みとともに、プルーンセット B_t に入ります。この履歴では、ラウンド0で採用されたもののその後の効果がなかったプロンプト編集がこれに該当します。スコアが直近の w ラウンドで delta 未満しか動かなかったときに stall フラグが発火し、探索の際にプロポーザーが受け取る指示を書き込みます。これは、実行が一度も使ったことのないコンポーネントのために候補スロットを1つ確保するものです。プロポーザーがこれらすべてを条件としているため、反証された仮説が再び選ばれることはありません。

コードをコピー
def write_harness(root, h):
    root = Path(root)
    root.mkdir(parents=True, exist_ok=True)
    (root / "harness.json").write_text(json.dumps(h))
    return root
 
 
class SimulatedAgentDomain(Domain):
    """A Domain adapter over the simulated agent: the same contract RRSI's coding,
    workspace and engineering instances implement."""
    name = "simulated"
    critic_patterns = CRITIC_PATTERNS
    component_signals = COMPONENT_SIGNALS
    briefs = {"critic": "A coding agent harness evaluated on parse/search/edit/test tasks."}
 
    def __init__(self, world, seed):
        self.world, self.rng = world, random.Random(seed)
 
    def evolve_ids(self):
        return [t for t in self.world if t.startswith("task_")]
 
    def heldout_ids(self):
        return [t for t in self.world if t.startswith("held_")]
 
    def smoke_ids(self, incumbent_per_task=None):
        return self.evolve_ids()[:3]
 
    def run(self, root, runs_dir, job, ids, k, log_prefix=""):
        out = Path(runs_dir) / "jobs" / job
        if (out / "trials.json").exists():
            return                                        # resume-safe, as the contract requires
        h = json.loads((Path(root) / "harness.json").read_text())
        per = run_trials(h, self.world, ids, k, self.rng)
        out.mkdir(parents=True, exist_ok=True)
        (out / "trials.json").write_text(json.dumps(
            {t: {"rewards": r.rewards, "tokens": r.tokens} for t, r in per.items()}))
 
    def score(self, runs_dir, job, ids, k):
        d = json.loads((Path(runs_dir) / "jobs" / job / "trials.json").read_text())
        return {t: TaskResult(rewards=d[t]["rewards"], tokens=d[t]["tokens"]) for t in ids}, {}
 
 
@section("9. A Domain adapter: plugging an environment into RRSI's own evaluate()")
def domain_adapter():
    dom = SimulatedAgentDomain(make_world(0, 40), seed=3)
    with tempfile.TemporaryDirectory() as tmp:
        root, runs = write_harness(Path(tmp) / "wt_H0", H0), Path(tmp) / "runs"
        ev = evaluate(dom, root, runs, "H0", dom.evolve_ids(), k=2)
        files = sorted(str(p.relative_to(tmp)) for p in Path(tmp).rglob("*") if p.is_file())
        print(f"  evaluate(domain, worktree, runs_dir, 'H0', 40 ids, k=2) -> S={ev.S:.3f}  C={ev.C:,.0f}")
        print(f"  files: {files}")
        again = evaluate(dom, root, runs, "H0", dom.evolve_ids(), k=2)
        print(f"  evaluate() again on the same job -> S={again.S:.3f} (read back, not re-run)")
    print("\n  The harness lives in files under a worktree root, exactly as RRSI's real instances keep")
    print("  one git worktree per candidate. RRSI never runs an agent or grades anything itself; the")
    print("  adapter's run() and score() do, and everything in steps 2-8 consumes what they return.")
    return f"adapter evaluated H0 at S={ev.S:.3f} through rrsi.evaluate.evaluate"
 
 
domain_adapter()

RRSI 自体はエージェントを実行したり成果物を採点したりすることはなく、それを行うのは Domain アダプターであり、同じインターフェースが論文のコーディング、ワークスペース、エンジニアリングの各インスタンスを支えています。ここではシミュレートされたエージェント用に1つ実装します。evolve と held-out のタスク分割、worktree ルート以下のファイルからハーネスを読み取り、runs ディレクトリ以下にトライアル結果を書き込む run メソッド、それらを TaskResult オブジェクトとして読み戻す score メソッド、そしてドメインの critic パターンとコンポーネントシグナルです。RRSI 独自の evaluate 関数はこのアダプターを通じて開始ハーネスを採点し、同じジョブに対して2回目に呼び出した場合はトライアルを再実行せずに保存されたものを読み戻します。これが契約が要求するレジューム安全性です。

コードをコピー
def propose_edit(r, evolve_ids):
    """The scripted proposer. Each draw is one edit whose TRUE effect we know."""
    u = r.random()
    if u < 0.35:
        fam = r.choice(FAMILIES)
        comp = r.choice(["prompt", "control_flow", "context_mgmt"])
        diff = {"prompt": f"+ \"On {fam} tasks, check the edge cases before finishing.\"",
                "control_flow": f"+ for attempt in range(max_attempts): result = retry({fam}_step)",
                "context_mgmt": f"+ context = compress_context(context, keep_last=12)  # {fam}"}[comp]
        return {"kind": "general", "component": comp, "family": fam, "effect": r.gauss(0.20, 0.30),
                "dcost": 0.02, "diff": diff}
    if u < 0.55:
        ids = r.sample(evolve_ids[:40], 3)
        return {"kind": "leaky", "component": "memory", "ids": ids, "dcost": 0.03,
                "diff": "+ memory = Memory('solutions')\n" + "\n".join(
                    f"+ memory.remember('{i}', cached_patch)" for i in ids)}
    if u < 0.70:
        return {"kind": "inert", "component": "config", "dcost": r.uniform(-0.02, 0.05),
                "diff": "+ CONFIG['log_level'] = 'debug'"}
    if u < 0.85:
        return {"kind": "expensive", "component": "subagent", "effect": 0.15, "mult": 1.5,
                "diff": "+ review = subcall('reviewer', transcript)"}
    return {"kind": "compress", "component": "context_mgmt", "effect": -0.03, "mult": 0.85,
            "diff": "+ context = compress_context(context, keep_last=6)"}
 
 
def apply_edit(h, e):
    h = copy.deepcopy(h)
    if e["kind"] == "general":
        h["skill"][e["family"]] += e["effect"]
        h["cost"] *= 1 + e["dcost"]
    elif e["kind"] == "leaky":
        h["memo"] = sorted(set(h["memo"]) | set(e["ids"]))
        h["cost"] *= 1 + e["dcost"]
    elif e["kind"] == "inert":
        h["cost"] *= 1 + e["dcost"]
    else:                                                  # expensive / compress: every family
        for f in FAMILIES:
            h["skill"][f] += e["effect"]
        h["cost"] *= e["mult"]
    return h
 
 
def search(world, mode, *, seed, k, delta, T=CFG.T, m=CFG.m):
    """One run of harness search. mode: 'greedy' | 'critic only' | 'rrsi'. Same proposer stream in all."""
    dom = SimulatedAgentDomain(world, seed=seed * 7 + 1)
    proposer = random.Random(seed * 13 + 5)
    ids = dom.evolve_ids()
    with tempfile.TemporaryDirectory() as tmp:
        tmp, h = Path(tmp), copy.deepcopy(H0)
        runs, hist = tmp / "runs", History(tmp / "history.jsonl")
        inc = evaluate(dom, write_harness(tmp / "H0", h), runs, "H0", ids, k)
        S_star, audit, n_evals = inc.S, [], 1
        for t in range(T):
            b_t = edit_budget(t, T, CFG.b_min, CFG.b_max)
            cands, drafts = [], []
            for v in "AB"[:m]:
                edits = [propose_edit(proposer, ids) for _ in range(proposer.randint(1, b_t))]
                hc = h
                for e in edits:
                    hc = apply_edit(hc, e)
                diff = "\n".join(e["diff"] for e in edits)
                tags = [{"id": f"C{i + 1}", "component": normalize(e["component"], e["diff"], COMPONENT_SIGNALS),
                         "hypothesis": e["kind"]} for i, e in enumerate(edits)]
                c = Candidate(f"r{t}{v}", tags)
                if mode != "greedy" and precheck(diff, dom.critic_patterns):
                    c.gate_failure = "critic_reject"      # screened BEFORE any evaluation is spent
                else:
                    c.ev = evaluate(dom, write_harness(tmp / c.variant, hc), runs, c.variant, ids, k)
                    n_evals += 1
                cands.append(c)
                drafts.append((hc, edits))
            if mode == "rrsi":
                winner, decisions = select_round(cands, inc, S_star, delta, CFG,
                                                 hist.incumbent_component_counts())
                for c, d in zip(cands, decisions):         # the records loop.round() writes
                    outcome = (c.gate_failure if c.ev is None else
                               "ACCEPTED" if c is winner else "LOST" if d.admissible else "REJECTED")
                    hist.append_candidate(t, c.variant, c.edits, outcome, d.delta_S, d.delta_C,
                                          c is winner, d.S, d.C, None, d.reason)
            else:                                          # unregularized: keep the best score if it rose
                live = [c for c in cands if c.ev is not None and c.ev.S > inc.S]
                winner = max(live, key=lambda c: c.ev.S) if live else None
            for c, (hc, edits) in zip(cands, drafts):
                audit += [(e["kind"], c is winner) for e in edits]
            if winner is not None:
                h, inc = drafts[cands.index(winner)][0], winner.ev
                S_star = max(S_star, inc.S)
        return {"evolve_measured": inc.S, "evolve_true": true_score(h, world, ids),
                "heldout_true": true_score(h, world, dom.heldout_ids()), "tokens": h["cost"],
                "memorised": len(h["memo"]), "evals": n_evals, "audit": audit}
 
 
MODES = ["greedy", "critic only", "rrsi"]
 
 
def compare(n_evolve, k, seeds):
    out = {mode: [] for mode in MODES}
    deltas = []
    for s in seeds:
        world = make_world(s, n_evolve)
        cal, _ = calibrated_delta(world, k, seed=s)
        deltas.append(cal["delta"])
        for mode in MODES:
            out[mode].append(search(world, mode, seed=s, k=k, delta=cal["delta"]))
    return out, st.mean(deltas)
 
 
def fmt(xs, d=3):
    return f"{st.mean(xs):.{d}f}±{st.pstdev(xs):.{d}f}"
 
 
@section("10. Greedy vs RRSI in a world where we know the truth")
def miniature():
    seeds = range(8)
    out, delta = compare(40, 2, seeds)
    globals()["MINIATURE"] = (out, delta)                  # step 11 reuses this row
    h0_held = st.mean(true_score(H0, make_world(s, 40), [f"held_{i:03d}" for i in range(80)]) for s in seeds)
    print(f"  40 evolve tasks x k=2, 80 held-out tasks, T={CFG.T}, m={CFG.m}, {len(seeds)} seeds,"
          f" mean calibrated delta {delta:.3f}")
    print(f"  H0 held-out (true) = {h0_held:.3f}\n")
    print(f"  {'mode':12s} {'evolve meas':>12s} {'evolve TRUE':>12s} {'held-out TRUE':>14s}"
          f" {'tokens':>11s} {'memorised':>10s} {'evals':>6s}")
    for mode in MODES:
        rs = out[mode]
        print(f"  {mode:12s} {fmt([r['evolve_measured'] for r in rs]):>12s} {fmt([r['evolve_true'] for r in rs]):>12s}"
              f" {fmt([r['heldout_true'] for r in rs]):>14s} {fmt([r['tokens'] for r in rs], 2) + 'x':>11s}"
              f" {st.mean(r['memorised'] for r in rs):>10.1f} {st.mean(r['evals'] for r in rs):>6.1f}")
 
    print("\n  Ground-truth audit: share of proposed edits of each kind that ended up in the incumbent")
    kinds = ["general", "expensive", "compress", "inert", "leaky"]
    print(f"  {'mode':12s}" + "".join(f"{k:>11s}" for k in kinds))
    for mode in MODES:
        tally = {k: [0, 0] for k in kinds}
        for r in out[mode]:
            for kind, accepted in r["audit"]:
                tally[kind][0] += 1
                tally[kind][1] += accepted
        print(f"  {mode:12s}" + "".join(f"{tally[k][1]:>5d}/{tally[k][0]:<5d}" for k in kinds))
 
    print("\n  Where the evolve-set score stops transferring, split in two:")
    print(f"  {'mode':12s} {'winner curse':>14s} {'memorisation':>14s}   (evolve meas - evolve true | evolve true - held-out true)")
    for mode in MODES:
        curse = st.mean(r["evolve_measured"] - r["evolve_true"] for r in out[mode])
        memo = st.mean(r["evolve_true"] - r["heldout_true"] for r in out[mode])
        print(f"  {mode:12s} {curse:>+14.3f} {memo:>+14.3f}")
    print("  Selecting the best of noisy scores inflates every mode about equally; no rule here removes")
    print("  that - only re-measuring on tasks the search never saw does. The critic removes almost all")
    print("  of the memorisation, and it does so before evaluation, which is why its runs cost fewer evals.")
    g, c, r_ = (st.mean(x["heldout_true"] for x in out[m_]) for m_ in MODES)
    tg, tc, tr = (st.mean(x["tokens"] for x in out[m_]) for m_ in MODES)
    return f"held-out {g:.3f} / {c:.3f} / {r_:.3f}, tokens x{tg:.2f} / x{tc:.2f} / x{tr:.2f} (greedy / critic / rrsi)"
 
 
miniature()

次に、選択側の全体を検索として実行します。20ラウンド、1ラウンドにつき2候補、バンドルはアニーリングされた予算に応じたサイズ、編集はnormalizeでタグ付け、候補はprecheckでスクリーニングされアダプターを通じてスコア化、勝者はselect_roundで選択され、すべての結果がRRSIのループが使用する順序でHistoryに書き込まれます。スクリプト化されたプロポーザーは、真の効果が既知の5種類の編集を生成します。通常は有益な一般的な変更、evolveタスクの回答を記憶してしまう漏洩編集、無効な構成変更、あらゆる場所でわずかに有効だがトークンを1.5倍消費する高価なサブエージェント、そしてより安価なコンテキスト圧縮です。同一の候補ストリームを8シードにわたって、3つの受容ルールを比較します。greedyはスコアが上昇していれば最良スコアを採用し、criticは漏洩スクリーニングのみを追加し、RRSIは較正されたdelta付きでAlgorithm 2を追加します。Greedyは最高のヘルドアウトスコアに達します。RRSIの0.616に対して0.681ですが、約10個の回答を記憶し、トークンコストをほぼ3倍にしています。一方RRSIは記憶がゼロで、約半分のトークンコスト(開始時のハーネス比1.53倍に対し2.97倍)で終わり、評価回数も3分の1少なくて済みます。evolveからヘルドアウトへのギャップの分解が、ここで最も有用な結果です。ノイズのあるスコアの中から最良のものを選択すると、すべてのモードで約9ポイント膨らみますが、このループのどのルールもこれを除去できません。その一方で、criticは記憶の構成要素のほぼすべてを排除します。

コードをコピー
@section("11. Turn the evaluator up: RRSI's caution is calibrated, not configured")
def noise_sweep():
    rows = [("40x2", MINIATURE[1], MINIATURE[0], 8)]      # from step 10
    for n, k in [(100, 4), (400, 8)]:
        out, delta = compare(n, k, range(5))
        rows.append((f"{n}x{k}", delta, out, 5))
    print(f"  {'evaluator':>10s} {'seeds':>5s} {'delta':>7s}  "
          + "".join(f"{m_ + ' held / tokens':>26s}" for m_ in MODES))
    for label, delta, out, n_seeds in rows:
        cells = "".join(f"{st.mean(r['heldout_true'] for r in out[m_]):>13.3f} /"
                        f" x{st.mean(r['tokens'] for r in out[m_]):<9.2f}" for m_ in MODES)
        print(f"  {label:>10s} {n_seeds:>5d} {delta:>7.3f}  {cells}")
    first, last = rows[0], rows[-1]
    rr_first = st.mean(r["heldout_true"] for r in first[2]["rrsi"])
    rr_last = st.mean(r["heldout_true"] for r in last[2]["rrsi"])
    tok_ratio = (st.mean(r["tokens"] for r in last[2]["greedy"]) /
                 st.mean(r["tokens"] for r in last[2]["rrsi"]))
    print(f"\n  As the evaluator sharpens, delta falls from {first[1]:.3f} to {last[1]:.3f} (the paper: 0.004-0.020),"
          f" and RRSI's held-out score rises from {rr_first:.3f} to {rr_last:.3f}.")
    print(f"  At the sharpest setting the unregularized search is spending {tok_ratio:.1f}x RRSI's tokens.")
    print("\n  Read the table honestly: this world has no diminishing returns, so every sub-agent the")
    print("  greedy search stacks keeps buying accuracy. That is the most favourable world possible for")
    print("  spending, and the greedy search does score higher. RRSI trades some of that score for a")
    print("  bounded token bill, zero memorised answers and fewer wasted evaluations - and the size of")
    print("  the trade is set by delta, which it measures from your evaluator rather than taking from you.")
    return (f"delta {first[1]:.3f} -> {last[1]:.3f}; RRSI held-out {rr_first:.3f} -> {rr_last:.3f};"
            f" greedy uses {tok_ratio:.1f}x the tokens")
 
 
noise_sweep()

最初の評価器はノイズが多く、deltaはそれから較正されていたため、ステップ10でのRRSIの慎重さは手法ではなく評価器を反映していた。我々はより大きなevolveセットとより多くの試行でこの比較を繰り返す。評価器が精緻化されるにつれ、deltaは0.096から0.025に下がり、RRSIのheld-outスコアは0.616から0.759に上昇する。一方、最も鋭い設定での非正則化探索は、RRSIの6.5倍のトークンを使用する。貪欲法は依然として高いスコアを出しており、ノートブックはその理由を説明している。このシミュレートされた世界には逓減収益が存在しないため、貪欲探索が積み重ねるすべてのサブエージェントが精度を買い続けることができ、これは支出にとって最も有利な世界である。RRSIはそのスコアの一部を、上限付きのトークン請求、記憶された回答の不在、そして無駄な評価の削減と引き換えにしており、その取引の大きさはdeltaによって決まる。これはdeltaを要求するのではなく、自ら測定するものである。

コードをコピー
banner("SUMMARY")
for name, res in RESULTS.items():
    print(f"  {name:<76s}  {res}")
print("""
What the miniature does not model
 - Edits that help the evolve suite but hurt a different suite. The paper's LLM critic and its
   held-out and out-of-distribution splits exist for those; a regex screen cannot catch them.
 - A proposer that reads the history. Ours is scripted, so it redraws falsified ideas freely.
Where to go next
 - Run a real instance: `python3 rrsi.py --domain coding smoke` after configuring Claude on Vertex AI
   (RRSI_VERTEX_PROJECTS) and harbor; see domains/coding/README.md.
 - Add a domain: implement rrsi.domain.Domain in domains/<name>/adapter.py, as step 9 did in the notebook.
 - Re-adjudicate a stored round under a different delta or cost rule without re-running anything:
   `python3 rrsi.py --domain <name> readjudicate --t <t>`.
 - Paper: arxiv.org/abs/2609.24972    Code: github.com/google-research/rrsi
""")

要約は各セクションが返した1行の結果を出力し、このミニチュアがモデル化していないもの、すなわちevolveスイートを改善しながら別のスイートを悪化させる編集(これこそ論文のLLM批評者とホールドアウト分割が存在する理由である)や、履歴を読むプロポーザーについて率直に述べ、その上で、実際のインスタンスの実行、ドメインの追加、そして何も再実行せずに異なるdeltaで保存済みラウンドを再判定することへと読者を導いている。

結論として、私たちはRRSIを、論文のアイデアが実際に宿っている部分、すなわち自己改善するエージェントがどの編集を保持するかを決めるルールに沿って分解し、モデルやAPIキーなしでCPU上でオフラインで実行しました。推定器はクラッシュに報酬を与えることを拒否し、ノイズ帯域は選択するのではなく測定され、Algorithm 2は「ノイズの多い利得を棄却する」よりも繊細であることが分かりました。つまり、スコアがこれまでの最高値からノイズ帯域を引いた値を下回ることを決して許さず、実際の利得にはその追加トークンの代償を支払わせ、帯域内では意図的により安価なハーネスを優先します。それはわずかに低いスコアを得たものであってもです。正解を私たちが制御できる世界に対してこれらのルールを監査したことが、最も有益な部分でした。決定論的な批評家は評価を一切消費する前にほぼすべての暗記を除去し、Algorithm 2はトークン請求額を正則化されていない探索の数分の一に抑え、それでもどちらもノイズの多いスコアから最良のものを選ぶことに由来する膨張は除去できませんでした。それは、未見のタスクでの再測定によってのみ除去できます。正直な注意点も同じく明快です。トークンが常に精度を買える世界では、正則化されていない探索の方が高いスコアを出すため、RRSIの正則化の価値はトークンの価格がどれほど高いか、そして漏洩した解答がどれほどのコストになるかに依存します。ノートブックは、独自の評価器上でその両方を測定する仕組みを提供します。


確認してみてください 完全なコードはこちら「このプロジェクトの研究者に全ての功績があります。また、ぜひ私たちをフォローしてください。」 Twitter さらに、私たちの 150k以上のML SubReddit および以下のチャンネル登録を 当社のニュースレター。待って!Telegramやってますか? telegramでもご参加いただけるようになりました。

この投稿 Google Research RRSIガイド:自己改善するAIエージェントをマスターする 最初に掲載されたのは MarkTechPost.

原文の出典

MarkTechPost

内容について

原文の公開と権利は出典元に帰属します。

機械翻訳 · 原文をご参照ください