A reweights the student with log-probability contrast; Z uses relative distance; R matches the feedback-conditioned teacher. The vertical axis is the student's probability of the correct token. Base teacher capability ρ moves the starting correct-token probability from ⅓ to 90%. Instruction following α controls feedback influence: the prior weight is s = 4(1 − α), and the clean posterior is (s · base + feedback) / (1 + s), with feedback 99% correct. At α = 0 the prior has four times the feedback weight; at α = 1 only feedback remains. These are toy-model proxies for capability and instruction following. Each round boosts one randomly selected wrong-token logit by Uniform(0, 1). A 32→128→3 network trains on 1,000 examples for 400 rounds, with 25 Adam updates per round, β = 10, and equal forward and reverse KL. The live curves use one paired seed; the maps average five seeds. Map outlines and values refer to the nearest measured cell; the dot tracks the continuous setting.