Live trainingPreparing the experiment
0.50
Uniform base90% correct
Interpolates the base teacher from uniform to 90 percent probability on the correct token.
0.00
WeakerStronger
Higher values give feedback more influence. This toy control runs from the strongest prior in the sweep to feedback alone.
Probability of the correct token over trainingA, Z and R learning curves.
Three-token toy model. Maps show five-seed means at round 400.
Toy setup

A reweights the student with log-probability contrast; Z uses relative distance; R matches the feedback-conditioned teacher. The vertical axis is the student's probability of the correct token. Base teacher capability ρ moves the starting correct-token probability from ⅓ to 90%. Instruction following α controls feedback influence: the prior weight is s = 4(1 − α), and the clean posterior is (s · base + feedback) / (1 + s), with feedback 99% correct. At α = 0 the prior has four times the feedback weight; at α = 1 only feedback remains. These are toy-model proxies for capability and instruction following. Each round boosts one randomly selected wrong-token logit by Uniform(0, 1). A 32→128→3 network trains on 1,000 examples for 400 rounds, with 25 Adam updates per round, β = 10, and equal forward and reverse KL. The live curves use one paired seed; the maps average five seeds. Map outlines and values refer to the nearest measured cell; the dot tracks the continuous setting.