What a session gives you
- A validated reward. Only calibrated, non-drift-flagged judges bound to the population you are training on. An unvalidated judge is refused with the reason, never downgraded.
- The certificate. Sensitivity, specificity and κ with intervals, the trust verdict, and what voids it — returned at creation, before the first rollout is scored.
- A ledger-true budget. The reward budget is held in your wallet at creation and every scoring call is metered; when the ledger crosses the ceiling the session refuses (
402 budget_exhausted). Your trainer never self-reports spend. - Masking, not zeros. A judge that cannot parse a verdict returns
grade: nullfor that item. Mask it out of the loss; treating it as 0 trains the policy to confuse the parser. - The anchor. Optionally, a second independent judge that holds the run when the reward diverges from it. See below.
Create a session
certificates[] and, if you asked for one, anchor. Sessions expire (ttlHours, default 72) and can be stopped with DELETE /v1/reward/sessions/{id}, which releases the hold.
Score a rollout group
step namespaces the billing ids, so re-sending the same step and request ids after a crash settles into the same ledger rows. Agentic sessions send steps[] (assistant and tool turns) instead of response; the server renders the trajectory through the same instrument your trace judge was calibrated on.
Adapters
Theerrorbar_reward Python package (in the errorbar repo under rl/errorbar_reward) exposes the same contract in three shapes:
ErrorbarRewardClient— create, read, stop, score, anchor, resume.trl_reward_func(client, session_id)— a TRLGRPOTrainerreward_funcscallable;Nonegrades pass through so TRL masks them.ErrorbarJudge— a Prime Intellect verifiers v1Judge; a masked item raises, matching verifiers’ rule that a malformed verdict must error the rollout, not score the model.
The mid-run anchor
In-loop statistics catch a policy collapsing onto a template. They do not catch the reward rising while true quality does not, because the policy found what the reward judge rewards rather than what the task wants. That is only visible from a measurement the policy never optimised against. The anchor is a second calibrated judge that must be a different instrument from every reward judge — a different model or prompt, not a clone with a new id — aligned to the reward bar, not drift-flagged, on the same unit. Every N optimizer steps your trainer generates one completion per prompt in a frozen prompt set and posts them:- Divergence: the reward rose at least 10 points over the last three checks while the anchor moved at most 2.
- Drop: the anchor sits at least 10 points below its best point so far.
/v1/reward/score answers 409 held, and the response names pinned_step, the step of the best anchor point — the checkpoint to ship. Points with fewer than 20 gradable prompts hold nothing; the response says underpowered instead. A hold is cleared by a person with POST /v1/reward/sessions/{id}/resume; it is never automatic.
Hold, never kill. A false hold costs an interruption you can reverse; a missed hack costs a shipped regression. Every verdict carries the anchor’s interval and n, so the false-hold rate is a number you can read rather than a promise.
Independence, stated precisely
Several tools publish grader agreement numbers, and judge alignment is a documented practice. What a reward session adds is a grader that is independent of the trainer, with intervals rather than point estimates, a refusal floor when underpowered, a certificate that expires and voids on change, and a re-check against an independent judge while the run is in flight.What you take away
A session is not only a reward function. It keeps every item it scored, and hands two things back. The graded trajectories.GET /v1/reward/sessions/{id}/export?format=rl|sft|pairwise returns JSONL: rl is one line per scored item (prompt, completion, reward, verdict, each judge’s grade, flags), the file you replay or re-weight under another algorithm; sft is the passing winners, strict verdict and reward at or above min_grade (default 0.9), in the fine-tuning prompts format; pairwise is chosen/rejected pairs from the same prompt within a step with a reward gap of at least 0.2, ready for DPO. The export scans at most 10,000 rows and says so in X-Errorbar-Export-Capped; narrow with step.
GET /v1/reward/sessions/{id}/environment returns a signed bundle: the reward spec, each judge’s calibration certificate, the deterministic assertions, the env tools you declared (names only), the tasks you stored on the session, the anchor and whether it is holding, and adapter snippets for TRL and verifiers. It is signed with the same key as eval evidence bundles, so POST /v1/verify checks it. ?format=tasks returns just tasks.jsonl, the shape our RL harness and OpenEnv-style loaders read. Pass tasks when you create the session and the whole environment round-trips.