At Alquist, we build robots for tasks in our partners’ physical environments. An LLM handles high-level planning: it interprets a request, chooses tools, and communicates with the person supervising the robot. Getting the task right requires following the rules of the environment, the expectations of the supervisor, and the requirements of the task itself.
As an illustrative example, consider a request to find an object. The planner might need to confirm the request before moving and respect restrictions on where the robot can navigate. It also needs to distinguish “not here” from “not found”: checking one area should not end the search while other permitted locations remain.
As we add tasks, the guidelines grow. So does the list of ways to violate them. Putting more instructions and exceptions into the system prompt becomes difficult to maintain. We want the planner to learn from corrections and carry that learning into future interactions. Some deployments also require a small model at runtime, so improving that model’s behavior matters more than simply calling a larger one.
This led us to study how to train a model from verbal feedback. A capable judge can review an attempt and explain what needs to change. The challenge is turning that explanation into an effective training signal for the model we deploy.
We turn verbal feedback into a distillation loss. We compare a teacher model’s predictions with and without feedback, then use that change to adjust the student’s own predictions. We also train on feedback-guided continuations, so the student can learn what comes after a corrected decision. With this approach, we substantially outperform the baselines in both knowledge acquisition and robot planning. The trained student acts without the feedback at inference time.
A score leaves the correction implicit
One way to train against a behavioral specification is to score each requirement, assign weights, and combine the results into a reward. The planner receives a higher reward when its actions satisfy more of the specification.
The hard part is assigning credit. A rollout is the model’s complete generated attempt, including its text and tool calls. If the planner searches one area correctly but skips confirmation and gives up too early, which choices in that rollout should training reinforce? Which should it discourage?
Common RL recipes such as GRPO apply the same reward-derived advantage to every token in an attempt. Across many attempts, the model can learn which decisions lead to higher rewards. But the reward itself leaves much of that work implicit. We can build more detailed reward functions, at the cost of more engineering for each clause and failure case.
A judge can instead provide a correction like this:
You started moving without confirming the request, then declared the object missing after checking only one area. Confirm the request with the supervisor before moving, and continue searching the remaining accessible areas before concluding that you can’t find the object.
This feedback identifies the skipped confirmation and explains how to continue after an unsuccessful search. A scalar reward compresses these instructions into a measure of how well the attempt went. Verbal feedback preserves them for the learner.
Turning feedback into a training signal
Distillation trains a student, the model we want to improve, using the predictions of a teacher. In the robot setting, the student is our planner. We use a frozen copy of the student’s initial model as the teacher, and a separate judge supplies verbal feedback.
At each position in a response, the student and teacher assign probabilities to possible next tokens. These probabilities provide a more detailed signal than a score for the whole response.
In on-policy distillation, the student generates the response. The teacher then scores possible next tokens at each prefix: the text generated up to that point. Training moves the student’s predictions toward the teacher’s at the situations the student actually encounters.
Feedback gives the teacher information the student did not have when it made the attempt. A judge reviews the rollout against the task criteria and writes a critique. We give that critique to the teacher, whose predictions now reflect the additional information. Figure 3 shows how this feedback enters the distillation loop.
The judge and teacher have different jobs. The judge can be a frontier model accessed through a text API. It can spend time reasoning about the attempt and the criteria before writing its feedback. We do not need its token probabilities. The teacher needs to expose probabilities over the student’s vocabulary and interpret the feedback at each prefix.
Designing that loss involves two choices: what target should the student learn from, and on which continuations should we evaluate it? We address each in turn.
1. Measure what the feedback changes
We first need a target for the loss. Matching the teacher’s entire distribution transfers many preferences at once. Some concern the correction. Others concern phrasing, formatting, or habits the teacher already had.
For the object-finding task, we want feedback about confirming a request to increase the likelihood of that interaction. Differences in unrelated wording should not become a reason to retrain the student.
We therefore score the same prefix twice with the same teacher. Both versions see the task and the student’s original attempt; one also sees the feedback. The difference tells us how the teacher responds to that feedback.
Learning from an imperfect teacher
We built a small experiment to see how the choice of training target affects what the student learns. There are just three possible answers, one correct. We simulate a teacher that receives useful feedback but also makes noisy changes to the probabilities of wrong answers.
The interactive Figure 4 trains three small neural networks with the same starting weights, examples, and training procedure. Only the target changes:
- Log contrast (A) reweights the student according to proportional changes in the teacher’s probabilities.
- Relative distance (Z, ours) reweights the student according to how far feedback moves each probability toward certainty or impossibility.
- Teacher matching (R) trains the student to copy the teacher’s probabilities after feedback.
The controls vary how much the teacher knows before feedback and how strongly it responds to feedback. The curves show the learning progress at your chosen setting; the maps show the final result across both controls. Darker green means a higher probability of the correct answer.

Static preview of a completed training run.
Relative distance stays green across the whole map. Its average final correct-answer probability remains near certainty at every tested setting. Log contrast works well when the teacher responds strongly to feedback, but can fall below chance when a capable teacher responds weakly. Teacher matching largely inherits the teacher’s confidence. Relative distance can learn from the correction and become more confident in the right answer than the teacher supplying it.
What should count as stronger evidence?
Suppose feedback changes two of the teacher’s probabilities:
- Correct token: 95% → 98%. A proportional increase of about 3.2%.
- Wrong token: 1% → 1.2%. A proportional increase of 20%.
Log contrast measures proportional growth, so it gives the wrong token the higher score.
Relative distance measures how much of the remaining gap the feedback closes. The correct token moved 60% of the way from its starting probability to certainty. The wrong token closed only about 0.2% of its own remaining gap.
Our confirmation score applies this idea to both increases toward certainty and decreases toward impossibility. For a candidate next token , write the teacher’s probability before feedback as and its probability with feedback as . For , the score is
An increase is divided by the probability left to gain; a decrease is divided by the starting probability. For the correct token above, this gives . The score stays between −1 and 1, with unchanged probabilities receiving zero.
Why this score? It turns out that’s the only ordering that fits the rules.
We treat a possible next token as a hypothesis, and the feedback as evidence for or against it. The teacher’s probability before feedback is the prior; its probability afterward is the posterior.
To compare those probabilities, we use four requirements from probabilistic confirmation:
- Use the before-and-after probabilities. The score should depend on those two values.
- More support should give a higher score. For a fixed starting probability, a higher probability after feedback should count as stronger confirmation.
- Measure decreases proportionally. When feedback lowers a probability, the score should depend on the fraction of its original probability that remains.
- Keep a hypothesis and its opposite consistent. If the feedback supports token A more strongly than token B, it should count against “anything except A” more strongly than “anything except B.”
Together, these requirements determine an ordering of the tokens, as established by Crupi and Tentori’s representation theorem. They leave the numerical scale free. The relative-distance score used above follows this ordering.
The complement requirement explains the distinction. Feedback strongly lowers the probability of “anything except the correct token,” from 5% to 2%. It barely changes “anything except the wrong token,” from 99% to 98.8%. Yet the log-probability contrast ranks both the wrong token and its complement higher than their counterparts. This violates the requirement that those rankings reverse.
Build the target around the student
We use these scores to reweight the student’s current next-token probabilities. If the student assigns probability to token , we construct the corrected target as
The parameter controls how strongly feedback changes the target.
In the numerical example above, both tokens have positive scores, and the correct token gets the larger boost. Log contrast reverses that ranking. Throughout training, we rebuild the target from the student’s current predictions. Learning can therefore continue even after the student becomes more confident than the teacher.
This construction has a useful property: if feedback leaves the teacher’s predictions unchanged, the target equals the student’s current distribution. There is no local distillation update at that prefix. Because the scores are bounded, the target’s change relative to the student is also bounded for a fixed reweighting strength. The resulting target reflects the feedback through the teacher’s interpretation.
2. Learn how to continue after the correction
The second choice is where to evaluate the loss. Even with a better target, training only on student rollouts can leave much of a critique unused.
Returning to the object-finding task, the judge asks the planner to confirm the request before issuing a navigation call and to keep searching after finding the first area empty. Two illustrative continuations show why the source of the rollout matters:
| Student attempt | Feedback-guided teacher attempt |
|---|---|
| Navigate immediately | Confirm the request with the supervisor |
| Search one area | Navigate and search the first area |
| Stop and report “not found” | Continue to the next accessible area |
The histories diverge at the first decision. On the student’s attempt, the teacher can favor confirmation before the navigation call. Every later prefix, however, contains a history in which the robot moved without confirming. Once the student ends the search, there are also no further search steps to learn from.
Figure 3 shows the loop with a student-generated attempt. To evaluate our loss, we allocate part of the rollout pool to the teacher conditioned on feedback. The teacher’s attempts can provide prefixes after confirmation and during a continued search. Both kinds of rollout train the student against the corrected target. This provides supervision on how to continue before the student can reliably generate the whole sequence itself.
An objective that covers both paths
We want the student to learn both from its own attempts and from continuations suggested by feedback. Here, is the probability that the current student generates a complete rollout . The corrected distribution, , gives the probability of that same rollout if we use our feedback-corrected next-token predictions at every step. Each rollout probability is the product of the corresponding token probabilities along that rollout.
KL divergence measures disagreement between two probability distributions. It is asymmetric: each direction averages over rollouts from the distribution on its left. Our ideal objective adds the two directions:
We aim to minimize this sum, also known as the Jeffreys divergence, which is zero when the student and corrected rollout distributions agree. The first term provides supervision at prefixes visited by the student. The second provides supervision at prefixes visited by the corrected policy, including continuations the student rarely generates.
We estimate both terms using the same pool of student and feedback-conditioned teacher rollouts. Pooling lets every rollout contribute to both parts of the loss, giving us more learning signal from the same rollout budget. Importance weights account for how each prefix was sampled. You can find the derivation and practical approximations in our paper.
Putting this together, here is how we train on a task:
- The student makes an attempt. It receives the task and generates a rollout without feedback.
- The judge reviews the attempt. It checks the rollout against the task’s requirements and writes feedback explaining what should change.
- The teacher generates additional attempts. It receives the task, the student’s attempt, and the judge’s feedback. We pool its feedback-guided rollouts with the original student rollout.
- We turn the feedback into training targets. At each prefix in the pool, we compare the teacher’s next-token predictions with and without feedback. The confirmation scores from this comparison reweight the student’s current predictions to form the corrected target.
- We update the student. We train its predictions toward these targets using both directions of the loss. Every rollout in the pool contributes to both terms.
Learning facts the teacher did not know
We gave the judge 20 fictional facts that neither the student nor the teacher knew. In Trivia Fantasy, each fact describes where an object is located, and each question has ten answer choices. All new factual information reaches the student through verbal feedback.
We trained Qwen3.5-2B and evaluated its recall without feedback. This simple experiment benchmarks knowledge transfer from the judge to the student. Success requires the student to acquire information that the teacher could not supply from its original knowledge alone.
We compared our method with two baselines, SDPO and W2S-OPD, adapted to learn from the judge’s feedback.
Following the robot’s behavioral requirements
We then tested whether feedback could help our robot’s planner follow requirements on new tasks. In EmbodiedEval, our internal benchmark, the planner carries out requests through tool calls in three-turn conversations with a simulated user.
We trained Qwen3.5-4B and evaluated it on 240 held-out tasks across 12 suites. A task passes when the evaluation judge detects no violation of its requirements. We report the mean pass rate across suites.
We reduced the mean failure rate by roughly a third compared with the strongest baseline.
Learning from the content of feedback
A useful correction can identify a mistake, explain a requirement, supply a missing fact, or describe the next attempt. Our method uses the teacher’s response to that information to construct a training target, and feedback-guided continuations to teach the student how to carry out the correction. The student can then use what it learned without receiving the critique at inference time.
For Alquist, the appeal is practical: the person supervising a robot can explain what should change in the language they already use. We want those explanations to become part of how the robot behaves on future tasks, so it can keep learning as its tasks and requirements change.
You can find the full derivation, proofs, benchmark details, and additional experiments in our paper.
