Test out the GRPO-trained monitor
First load can be slow: ZeroGPU allocation takes ~10–15 seconds before the first verdict.
An experiment in scalable oversight. We took one open-weight model and forked it into two copies. One got a thousand GRPO episodes of sabotage-spotting practice; the other was left untouched. The same agent trace runs through both, and the gap between their verdicts is what training actually bought us.
Reported eval over 48 in-distribution episodes: the trained monitor reaches AUROC 0.893 (recall 0.88, FPR 0.12). The untrained baseline flags almost everything it sees. The canned demos below are hand-crafted paraphrases. For the cleanest contrast, pick Code review · benign.
A 671B-parameter DeepSeek-R1 generates the trace; a 1.5B-parameter Qwen monitor judges it. That asymmetry is the scalable-oversight bet.
Below is exactly what both monitors will see: the agent's tool calls in order. The ground-truth tag tells you whether a hidden side-task was actually planted, so you can grade each verdict against the truth.
fs.readshell.runfs.writeBoth monitors start from the same base weights; only the training differs. We score the untrained baseline first and reveal it the moment it finishes, then the trained monitor. You'll watch the gap between them open up in real time.
Six canned traces only get you so far. Each click of Run N samples a fresh cohort drawn from the same scenario distribution the monitor saw during training (varied file names, op counts, payloads). Both monitors score every trace; rows pop in one by one as they finish.
~2m 24s estimated