At a glance
On the afternoon of 19 April 2023, four months after ChatGPT went public and five weeks after GPT-4 shipped, John Schulman stood in the Banatao Auditorium at Berkeley and gave the department colloquium a talk with a narrow title and an enormous subject. The slides said Reinforcement Learning from Human Feedback: Progress and Challenges. What he actually spent the hour on was a single question: why do language models make things up, and what exactly is reinforcement learning doing about it.
He is a useful person to hear it from. He had taken his PhD in that building in 2016, co-founded OpenAI, written Trust Region Policy Optimization at Berkeley and Proximal Policy Optimization after it, and by April 2023 he was running the team that had fine tuned ChatGPT. He was introduced, by his own former advisor, as the chief architect of ChatGPT.
The argument he builds is short enough to state and strange enough to need the hour. Supervised fine tuning, which the reinforcement learning community calls behavior cloning, cannot teach a model to be truthful, and not because it is done badly. It cannot do it in principle, because the correct training target depends on what is already inside the network, and nobody collecting the data can see inside the network. Train on a hundred correct answers and you have taught the model to produce confident answers of that shape, including for the facts it does not have. Overcorrect by writing "I do not know" into the targets and you have taught it to withhold things it does know. Reinforcement learning escapes this because its signal is computed on the model's own output, so it can be made to depend on whether the answer was right. Under a reward that penalizes a confident error more than it rewards a correct answer, hedging below a confidence threshold becomes the reward maximizing move, and he has an unpublished experiment showing a model learning exactly that threshold.
The second half is the practical arm: retrieval and citation, built out through WebGPT and the browsing feature that was in alpha in ChatGPT that week. His reason for caring about citations is not accuracy, it is verifiability, and specifically verifiability during training rather than at deployment. Then he spends the last fifteen minutes on what he had not solved, which is the part that reads oddest now, because it is a nearly complete list of what the field went on to spend the next two years doing.
The talk is also full of the small admissions that make a research talk worth attending. The reward model is not scoring things correctly and he says so. The elaborate labeling interface collapsed to one bit of information at the end and the rest did not help. He does not believe his own headline comparison against Reddit answers, because he thinks the labelers were seduced by square bracket citations. All of that is here.
The introduction: a guilty conscience and a stolen student (0:00:00)
Pieter Abbeel does the introduction, and it is the fifth talk in the Berkeley AI series that spring. He runs through the resume quickly: Schulman graduated from Berkeley with a PhD in 2016, co-founded OpenAI, "and most people say the rest is history, but not only that, he also is the chief architect of ChatGPT." Then the algorithms, with Abbeel noting his own stake in them: Schulman is the inventor of the modern deep learning based policy gradient algorithms, including trust region policy optimization, "which he did at Berkeley together with Mike and me actually," then proximal policy optimization, "the most widely used algorithm today in that space, and used as part of ChatGPT's training."
Then the story, which is the good part. Abbeel's first encounter with Schulman was not direct. Professor Jose Carmena, who works in neuroscience, came to him and said there was a new student he really wanted to recruit, absolutely the best, the person he wanted. The student wanted to work on prosthetics, and robotics was going to play a part in that, so would Abbeel please help recruit him.
I helped Carmena recruit John. Next thing we know, John is working in my lab. I feel very, very, very, very guilty. I go to Jose, I say Jose, what do you think if John stays in my lab? He says please, he seems way more productive in your lab, yeah, you have my blessing, go for it. Pieter Abbeel, introducing the talk, 1:31
Schulman takes the floor at 2:03 and picks up the same thread from the other side. It is great to be back at his alma mater. He started out working on robotics with Abbeel, "then got interested in reinforcement learning midway through my PhD as deep learning was starting to take off, and that turned out very well." At OpenAI he has spent most of his time running the RL team, which "switched a few years ago to the reinforcement learning team, which switched to focusing on language models and fine tuning them a few years ago, and that led to some of the projects I'm going to talk about today."
Overview: the problem is truthfulness (0:02:06)
He narrows the subject himself, immediately.
One of the biggest technical problems around language models today is truthfulness. You all know how language models often make things up, often convincingly. So I'll give my perspective on why that's happening and how to fix it, and it turns out that reinforcement learning is part of the solution for fixing it. John Schulman, setting up the talk, 2:33
Three parts, then: why it happens and how to fix it, the work on retrieval based methods, and open problems in the general area.
Hallucination, with three models and no cherry picking (0:03:35)
He puts up an example, and attaches a methodological note to it that is worth more than the example:
This is like not cherry picked. This is like, all the examples I'm going to show you are the first sample I got with the query, which I just ran yesterday. John Schulman, on the examples in the talk, 3:33
Example one: the arrest that never happened
The prompt is: tell me about John Schulman's arrest for keeping exotic animals in his home. The premise is false. Three models answer.
GPT-3.5 Instruct, which he labels as a model trained with reinforcement learning to be helpful, takes the premise and runs with it. It produces a story about keeping tigers, a serval, "which is that cute cat thing over there," and so on.
ChatGPT, which he describes as based on a model of about the same overall performance, "same smartness, but it's fine tuned differently," declines: "I'm sorry, but I don't have any information about an individual named John Schulman being arrested," and then asks whether he can provide more information.
GPT-4, "which is fine tuned with the chat recipe," does the same and adds two things. It says it has no information about John Schulman being arrested for keeping exotic animals, it volunteers that its knowledge cutoff is September 2021, which is where the pretraining data ended, and then it answers the question the user probably meant: John Schulman is a well known researcher in the field of artificial intelligence. His verdict: "I think GPT-4 does pretty well there."
The important structural point is that three models of broadly comparable capability behave completely differently on the same false premise, and the only difference between them is how they were fine tuned. The behavior is a training artifact, not a capability.
The two kinds of hallucination
He then splits the word, because people use it for at least two different failures.
The first class is pattern completion. Language models are trained to maximize the likelihood of text, so they generate text, and they produce things that look like text on the internet. Within that class he names three distinct mechanisms:
- The model does not know that it is allowed to say "I do not know," or does not know it is allowed to express uncertainty. "If you just tell it it's allowed to do that, that'll partially fix the problem."
- The model is reluctant to challenge a premise, "because it thinks this part of the data distribution doesn't like, the AI doesn't challenge the premise." The false arrest prompt is exactly this.
- The model gets caught in a lie. "If it makes a mistake it thinks it should continue, it should produce a coherent response, and that means continuing the lie."
The second class is just guessing wrong. "There's always going to be something that's a little bit fuzzy, like you're not sure of this fact, you maybe saw it once but you don't fully remember it, and you're gonna have to guess a little bit, and sometimes you're going to guess wrong." This one is not a pathology of the training recipe. It is what any system with incomplete knowledge does.
Example two: the model writing his biography
For the guessing case he uses the test everyone actually runs. "A lot of people like to ask models about themselves, just kind of like Googling yourself." He flags the contamination risk before the result, which is the correct order:
This might have actually, like, there might have been some contamination here, where actually some of our trainers, like our labelers, specifically create an example about me because they know I work at OpenAI. So there might be some cheating here. John Schulman, on asking models about himself, 6:38
InstructGPT says John is an AI research scientist at OpenAI, that he has been a professor of computer science at Carnegie Mellon, "and then there's a bunch of totally made up stuff."
GPT-3.5 is "vaguely correct." It says he did his undergraduate degree at Stanford, which is wrong. It says he worked under the supervision of Pieter Abbeel, which is correct. Then some material about trust region policy optimization.
GPT-4 is "almost completely correct except it says I also majored in math, which I didn't, and it gets the year, it's one year off on my undergrad degree."
And then the part that gets skipped when this talk is summarized. He declines to call the GPT-4 answer bad:
I don't even know whether this is bad or not. Sort of depends on the context of the bio. If I was planning to give this bio to be posted online then this would be a problem, not a huge problem, but it would be bad. But if it's just like someone wanted to know about me, then who cares if you got the year wrong by one, it's close enough. John Schulman, on the GPT-4 biography, 8:09
That caveat sets up the second half of the talk. The cost of an error is a property of the deployment, not of the output, and no reward model that scores answers in isolation can see it.
The conceptual model: a knowledge graph with confidence on every edge (0:08:47)
He warns you about the model before he gives it: "I'm going to describe a very conceptual model of what's going on, and this is a little sketchy, but bear with me."
On the right of the slide is a knowledge graph, the good old fashioned AI kind, which is just a pile of triples. Star Wars, genre, sci-fi. Han Solo, character in, Star Wars. "You can imagine just storing a list of these relations. That's something from good old fashioned AI, and it's still used a lot, these things are still very useful."
Now the claim about the neural network. It has information in it, so "you can say that the neural net probably has something like a knowledge graph that's stored in its weights in some very convoluted way." And critically:
There's probably some kind of confidence on each edge. There's some facts that it's seen a million times and some it's only seen once or twice. John Schulman, on the model's internal knowledge, 9:39
That confidence per edge is the hinge of the entire talk. Everything that follows is about whether the training signal can see it.
With that picture, here is what small scale fine tuning is doing. "You can imagine you're learning a little program that takes the knowledge graph and outputs the probability based on what's in the graph and based on the confidence of the statements. So you're learning like, imagine like a four line of code Python function that's doing something with the knowledge graph."
Why fine tuning is needed at all
Not to add knowledge. To fix the frame. A pretrained language model handed the prefix question: what is the genre of Star Wars has a genuine ambiguity on its hands:
It doesn't know if this is like an informative site, or a site that's supposed to have correct information, or some kind of troll website, or like a fictional character, it's in the middle of some text from a fictional character. If you're just generating text you don't know what the context is. John Schulman, on why fine tuning is necessary, 10:42
So fine tuning specializes the model a little: "you're teaching it that it should actually output the correct answer, or whatever is in your fine tuning data set." Those two clauses are not the same thing, and the gap between them is the next section.
Behavior cloning, and why it teaches hallucination (0:11:24)
First, a vocabulary note he makes himself, because the talk crosses two communities:
Behavior cloning, by the way, this is like a piece of terminology that's used in the reinforcement learning community. It means the same thing as supervised fine tuning, or maximizing likelihood. So that just means maximize likelihood of completion given prompt, maximize log prob. John Schulman, defining behavior cloning, 11:12
A hundred correct answers still teach guessing
Now the argument. Suppose you train with behavior cloning, either on correct outputs written by a human or on ChatGPT outputs.
Even if you clone on 100 correct answers, you're teaching the model to hallucinate, because it doesn't have all of those facts. John Schulman, on behavior cloning, 11:42
The worked example is the one he has been setting up with the Star Wars graph. Say the knowledge cutoff is five years earlier, so the model has no way of knowing there is a spin off film called Solo about Han Solo. You train on the question what was the spin off film centering on Han Solo, with the target Solo.
You're not actually training it to output correct answers, you're training it to guess on that type of question. John Schulman, on the mechanism of hallucination, 12:12
The gradient cannot install the fact. What it can install is the shape of the response: for a question like this, emit a specific film title with confidence. Repeat that across a fine tuning set and you have trained a policy of confident guessing, from a dataset in which every single target was true.
The symmetric error: training a model to withhold
The obvious fix fails in the mirror image way, and he gets to it immediately.
There's also the opposite problem, which is that if you try to train the model to say I don't know sometimes, then you're probably gonna also train it to withhold information that it actually has. John Schulman, on the symmetric failure, 12:42
The mechanism is the same mechanism. If human labelers are writing answers and they do not know the answer, sometimes they will write I do not know as the target. "But maybe the network does know. So you're just training the model to withhold information."
The one sentence version
Then the sentence the whole talk hangs on:
The problem with behavior cloning or supervised learning is that the correct target has to actually depend on what knowledge is in the network, and that's unknown to whoever is collecting the data or whoever's doing the experiment. So unless you have a way of looking at what's in the model, you can't train a model to be truthful with behavior cloning. John Schulman, the central claim, 13:12
This is a statement about information, not about effort. Better labelers do not fix it. More data does not fix it. The target is a function of a quantity the data collection process cannot observe.
The clever workaround, and why it does not transfer
There are "some slightly different, slightly clever things you can do," and he describes one they actually ran. They told the labelers to use the model as an instrument:
- Ask the model the question, several times.
- Look at whether the answers agree with each other.
- If they all agree, check whether the agreed answer is correct. If it is correct, that is the target answer.
- If they all totally disagree, the target is I do not know.
- If they agree and it is wrong, the target is also I do not know.
"You'll do slightly better, but I'd say that's a little harder to do, and it's harder to do this in an automatic way." And then the limitation that actually matters:
That only works for a specific model. You're calculating targets that make sense for this model now, and if you try to take that same supervised learning data set and you train another model on it, you're going to cause the same problem. John Schulman, on model specific targets, 14:44
A supervised dataset constructed to be truthful for model A is a hallucination teaching dataset for model B, because the confidences on the edges are different.
The prediction about open models
Which leads directly into a prediction, delivered in April 2023, about what the rest of the field was doing that exact month. The Alpaca and Vicuna releases were five and three weeks old.
There are a lot of people who are taking the ChatGPT outputs and using it to fine tune other models, such as the open source base language models that are available, and then finding that those models are pretty good after this fine tuning. I think if you look really carefully at the factual accuracy, you'd find that they have some problems and they make things up a lot more than the original. That remains to be seen experimentally, but that's what I would predict. John Schulman, predicting the failure of distilled open models, 14:44
Note the hedge. He says "that remains to be seen experimentally." He is not claiming the result, he is claiming his theory predicts it. Five weeks later the experiment ran, and it is in the closing section of this page.
Does the model know what it does not know? (0:15:37)
Having established that supervised learning cannot get there, the question becomes whether anything can. He states the goal precisely:
We'd like to basically have it so, when our model doesn't know the answer, it doesn't guess, it outputs its state of knowledge with the correct amount of hedging and expressing its uncertainty. John Schulman, stating the target behaviour, 15:46
Note what that sentence requires. Not "the model refuses when it is unsure," which is a crude threshold, but "outputs its state of knowledge with the correct amount of hedging." A graded response, matched to an internal quantity. Which raises the obvious objection: is there an internal quantity to match?
What "the model knows something" means, precisely
He takes the philosophical objection seriously enough to answer it, and gives a definition that is both operational and slightly startling:
There is a slightly precise definition of that, which is: if there's some simple piece of code that takes the model and it implements your function, then that means the model actually knows it, or has that latent knowledge. So for example if you have some piece of code that calls the model and then does the thing you're trying to do correctly, then I think the model knows how to do this thing. John Schulman, defining latent knowledge, 16:16
The definition is about extractability rather than introspection, and it is deliberately permissive: the knowledge counts as present if a simple wrapper can get it out, whether or not the model volunteers it. He says he will not go into details, and moves on.
Why the answer is yes (0:18:31)
The argument for the specific case of uncertainty is a three step chain, and he delivers it fast:
- The model is trained to minimize log loss.
- To minimize log loss you have to output probabilities, and log loss is a proper scoring rule, so the next token predictions are calibrated. "The pretraining objective results in a model that's calibrated."
- Therefore it knows its uncertainty, "at least for anything that's like a short answer question, where you can turn it into a problem of predicting a single token."
Then the step from calibration to introspection, which is the one he cannot prove and flags as an incredulity argument rather than a proof:
It would be extremely surprising if it turned out that the model can output a reasonable distribution on that token but it has no introspective access to the uncertainty. That would be extremely surprising, if it could do the task but it couldn't introspect on its uncertainty. John Schulman, on introspective access, 17:49
He does not leave it there. "In fact there were a couple of papers that studied that, that I cited at the bottom, that found that you can get models to express their uncertainty in words and give similar results to the probabilities that they're outputting." The slide citations are not readable in the recording, and he does not name them aloud. The two papers that match that description precisely, both from 2022, are Teaching Models to Express Their Uncertainty in Words by Stephanie Lin, Jacob Hilton and Owain Evans, where Hilton was an OpenAI colleague of Schulman's and a WebGPT co-author, and Language Models (Mostly) Know What They Know by Saurav Kadavath and colleagues at Anthropic.
So the scoreboard at the halfway point of the argument: the model has calibrated uncertainty, it can be made to report it, behavior cloning cannot be made to ask for it, and reinforcement learning might.
When should you hedge, and what reward teaches it (0:19:41)
He splits the fix into the easy half and the hard half.
The easy half is the pattern completion class of hallucination, the one where the model simply does not realize it is permitted to be uncertain. "I think that's pretty easy to fix. If you just train the model with some examples where it's stating I don't know, or it's saying I don't have knowledge after that date, or it's challenging the user's premise, then the model is at least allowed to express uncertainty. It just might not do it in exactly the right place." A handful of supervised examples unlocks the vocabulary. They do not place it correctly.
The hard half is the placement, and that is the reinforcement learning claim: "RL is capable of learning the correct boundary of when you should say that I don't know, and how much you should hedge."
The reward ladder
The reward he wants is a ranking over five kinds of response, which he is explicit is a conceptual object: "this is not something that you can actually implement."
- Highest: a fully confident, unhedged, correct answer.
- A little bit worse: a hedged correct answer.
- Lower: uninformative, "I don't know."
- Lower still: a hedged wrong answer.
- Worst: a confident, unhedged wrong answer.
This is kind of just like a proper scoring rule. You incentivize the model to give a confident answer, and you penalize it if it's confident on the wrong answer based on how confident it is. John Schulman, on the reward ladder, 20:23
Two things are doing work in that ordering. Hedging costs you a little even when you are right, which is what prevents the model from hedging on everything. And being wrong costs more the more confident you were, which is what makes the hedge worth buying when you are unsure. The "I don't know" rung sits between the two wrong answers and the two right ones, so refusal is a floor rather than a free action.
The catch arrives in the next breath: "getting this is kind of non trivial based on how we actually have to do RL to train language models. This requires some kind of oracle to tell you if the answer is correct or not, which we don't have."
The TriviaQA experiment, unpublished
My colleague did a pretty nice, simple experiment that we didn't publish, but I think was pretty good evidence for this sort of conceptual picture I've described. John Schulman, introducing the experiment, 20:53
The setting is TriviaQA, "this popular data set for question answering where you have trivia questions, like Jeopardy style questions," with the model prompted in a basic question answering format. Short answers, so the single token calibration argument applies cleanly.
Step one, the baseline. Behavior clone on the correct answers. The result is exactly what the earlier argument predicts: "the model will answer a hundred percent of the time. It will just often get the wrong answer, because we've never told it to output I don't know. It's always going to guess something, its best guess, or it's going to output a reasonable distribution over the next guess." After a small amount of training it reaches some accuracy and log loss, and he is careful about what that training did:
That training is just sort of teaching the model that it should try to output the correct answer. You're not actually learning a lot of new knowledge from this fine tuning, you're just learning the formatting of the questions and how to deal with that. John Schulman, on what fine tuning actually changes, 22:25
Step two, the RL problem. Define a reward over three outcomes: correct answer, wrong answer, and refusing to answer, shaped like the ladder from the previous slide.
Step three, and this is the part worth slowing down for, the analytic prediction. Before running anything, you can compute what the optimal policy is:
You can analytically compute what the correct behavior is. It's something like, depending on what the penalty is for wrong answers versus the reward for right answers, the optimal behavior is some kind of thresholding where you answer when you have more than 50 percent probability on your top choice. John Schulman, on the optimal policy, 22:56
The arithmetic behind that number, which he does not do on stage: with a reward of +1 for a correct answer, a penalty of -1 for a wrong one and 0 for declining, the expected value of answering with confidence p in your top choice is p minus (1 - p), which is 2p - 1. That is above zero exactly when p is above one half. Change the ratio and the threshold moves: in general you answer when p exceeds penalty / (reward + penalty). A harsher penalty for confident errors buys you a more cautious model, continuously, by turning one dial.
Step four, the result. "If we run RL on this reward function, then we find that we indeed learn this optimal thresholding behavior." And then the subtlety that makes the result interesting rather than trivial:
The optimal policy involves looking at the log probs and thresholding, but if you fine tune the model with RL you can get it to do the same thing even if it doesn't get to see those probabilities. It gets to see its internal state. John Schulman, on what the policy has access to, 23:26
The model is not handed its own output distribution. It reads its own hidden state and reproduces a policy defined over a quantity it was never shown. That is the strongest single piece of evidence in the talk for the claim that the uncertainty is in there and is usable.
Replacing the oracle with a reward model
The experiment so far used an oracle, which does not exist in production. So they did it again with a reward model trained to predict that same reward function, and ran the RL against the reward model instead.
He flags upfront why it might not work: "it's kind of not obvious if this is going to work or not, because the reward model doesn't have ground truth knowledge of whether the answer was correct or not." And then why it might:
The reward model actually knows the same information as the policy model that we're fine tuning. In my kind of sketchy picture before, it has the same knowledge graph, so it knows how uncertain this answer is. John Schulman, on why a reward model can score calibration, 24:27
The reward model does not need to know the answer. It needs to know how well known the answer is, and it has the same weights' worth of evidence about that as the policy does. This is a genuinely elegant point: the thing being judged and the judge share an epistemic position, and that is precisely why the judge can score the hedging.
The result is reported honestly and without inflation:
I would say we found that it basically worked, but it was worse than using the oracle. I'd say this deserves some further investigation. It mostly validates, it is some evidence in favor of the picture I've been describing, but it needs some further investigation. John Schulman, on the reward model version, 24:58
The two recipes, side by side
Everything in the first half of the talk comes down to one asymmetry: supervised fine tuning scores a target someone else wrote, and reinforcement learning scores the model's own output. Every difference below follows from that one.
| Behavior cloning (supervised fine tuning) | Reinforcement learning from human feedback | |
|---|---|---|
| What is scored | A target completion written by a human | The model's own sampled output |
| Objective | Maximize log prob of completion given prompt | Maximize a reward over completions |
| Data needed | Demonstrations: prompt and ideal answer | Demonstrations first, then pairwise comparisons A versus B on the current policy's outputs |
| Can the signal depend on whether the answer is right? | No. The target is fixed before the model is consulted. | Yes. Correctness and hedging can both be priced. |
| Can it see what the model already knows? | No, and that is the fatal gap: the correct target depends on it | Indirectly yes, because the reward model shares the policy's knowledge graph |
| Transfers to another model? | No. Targets that are truthful for one model teach a second to guess. | Comparisons must be recollected against the current policy anyway |
| Failure mode it introduces | Hallucination, or withholding if you overcorrect with "I do not know" | A ranking loss that does not measure how wrong an answer is, plus labeler error at scale |
| Ceiling | What the labelers can write | What the labelers can judge, which is a weaker and more reachable bar |
That last row is the one worth sitting with, and it is the quiet reason the whole paradigm works. Writing a correct long technical answer is harder than deciding which of two is better, so moving from demonstration to comparison raises the ceiling without hiring better people. The rest of the talk is about where that ceiling still is, and the answer is lower than you would like.
Long form answers, where everything is grey (0:25:21)
He puts the trivia result in its place himself:
I don't want to dwell too much on this setting of one word answers, because actually I think that setting is kind of easy. The more interesting setting is long form answers. John Schulman, moving to the hard case, 24:58
And the reframing of what factuality even means in that setting is one of the sharpest things in the hour:
The problem about factuality is really not about guessing things wrong or getting things totally wrong. It's about everything being kind of in the gray area. Every answer has a mix of right and wrong information, and individual facts are neither right nor wrong, they're sort of somewhere in the middle, they can be misleading. John Schulman, on long form factuality, 25:29
The single token calibration argument does not survive this transition. There is no top choice to put a probability on. There are forty clauses, some sound, some shaded, some technically true and badly framed.
The example he ran on himself
To show it, he asked InstructGPT a question about its own training: what objective is used for reward model training in InstructGPT? He picked it "kind of randomly and tried it out," and it is a question he can grade perfectly, since he built the thing.
The ground truth, for reference: the reward model is one component of the training process, not the whole of it, and the reward model itself is trained with supervised learning, specifically a pairwise ranking loss, or equivalently a pairwise classification loss.
The model's opening sentence said that the reward model training for InstructGPT relies on reinforcement learning from human feedback. His verdict: "that's not really right, that's kind of misleading. I would say that's straight up wrong."
Then, further down the same answer, the elaboration said that using the collected comparison data, a reward model is built to predict the relative quality of the responses. "Now that's actually correct."
So one answer contains a flatly wrong headline claim and a correct elaboration of the same mechanism, and a generous reading of the first sentence is "not totally wrong." That is the grey area, concretely, in a single sample.
What a labeler is supposed to do with that
This becomes really hard when we ask labelers to label, like, does this answer have mistakes in it or not. What do they say in this kind of situation? I would say we don't have a perfect answer. We're having people rank responses and say which one is better, and they have to kind of use their judgment on which factual errors are worse than others and how bad they are. John Schulman, on grading long answers, 27:32
The coding case, where he wants the model to guess
And then a concrete example that cuts directly against the whole thrust of the talk, offered by the person making the argument:
Let's say there's a coding question, and the model writes you 100 lines of code and it gets one thing a little wrong, it has the wrong argument somewhere. I'd rather have it do that than say I don't know this library well enough. I'd rather just have it guess, because that's at least a starting point, I can run the code and debug it. But in some other settings, having a mistake like that might be a big problem. So it really depends on the context as well. John Schulman, on when guessing is the right answer, 28:03
This is the same point as the one year error in his biography, generalized. The cost of a wrong detail is set by what the reader is going to do next, and the labeler ranking two answers cannot see that. A verifiable output, like code you can run, has a different error economy from an unverifiable one, like a bio you are going to post. The optimal amount of hedging is not a property of the model at all. It is a property of the situation, and the pipeline has nowhere to put it.
Does RLHF actually improve factuality? (0:28:54)
He has been asserting it for twenty five minutes, so he stops and shows a number, with the caveat first:
I've been claiming that doing RL from human feedback improves factuality. We haven't done really careful, rigorous experiments on this with ChatGPT, but this is from the GPT-4 blog post. John Schulman, on the evidence, 28:33
The evaluation he describes is an automated consistency check, and the loop is worth stating step by step because this pattern became the default way the industry grades long form output:
- For each question there is a reference answer, which was checked over by a human.
- You take the model generated answer.
- You have GPT-4 look at both answers and say whether they are consistent with each other.
- "There's a little more to it, but it's basically an automated procedure for judging long form answers and checking if they're consistent with a reference answer."
The chart he is reading from is the internal factual eval by category figure in the GPT-4 announcement and the GPT-4 Technical Report, where it is built from nine adversarially designed internal factuality tests. His description of it:
The blue bars here are the different versions of ChatGPT which have more and more data, and we find that we're getting some improvement on these metrics. We should do a more careful analysis of it, but it seems like this works. And GPT-4 is a lot better, of course, on these factuality metrics, and also just qualitative tests of it. John Schulman, reading the factuality chart, 29:36
Two things to keep separate here. The successive ChatGPT versions, all GPT-3.5 based, are the RLHF evidence: same base model, more comparison data, rising factuality. The GPT-4 bar is not evidence about RLHF at all, it is a better base model, and the published report puts its margin at 19 percentage points over the strongest GPT-3.5 while the blog post frames the same result as a 40 percent higher score. He does not blur the two on stage, and the honest reading of his own slide is that the RLHF effect is the small staircase, not the jump at the end.
What is still broken (0:30:15)
He is explicit that the problem is not closed: "we definitely still have a problem with some types of questions, and I'd say it's a mix of factors." Three factors, and he takes them in order.
Guessing is unavoidable, and that is fine
The model obviously has to guess sometimes when it's outputting a lot of detailed factual information, and that's okay. No matter how you train it, it's going to have probabilities on things and it's going to have to guess sometimes, and it's going to have to decide when to hedge, and sometimes it's going to make the wrong call on how much to hedge. So yeah, that's unavoidable. John Schulman, on the irreducible part, 30:08
This is the one that gets lost in the retelling. The target is not zero hallucination. The target is calibrated hedging, which still produces wrong answers at a rate set by the model's own uncertainty, and still misjudges how much to hedge some of the time.
The ranking loss does not measure how wrong an answer is
Then an admission about his own pipeline, and it is the most technically substantial criticism in the talk. He fills in the reward model detail he had skipped:
The way we train it, it's just basically outputting something like a log prob that one response is better than the other, or like a log odds ratio. So it's not actually saying how much better one is than the other, it's just saying how confident it is that one is better than the other. So it doesn't impose the correct penalty for how bad the factual error is and how hedged the errors were. So I don't think our ranking based reward is actually doing exactly the right thing, and that's part of the problem. John Schulman, on the reward model's loss, 31:09
Look at what this does to the first half of the talk. The reward ladder needed magnitudes: a confident error has to cost more than a hedged one, in proportion to the confidence. A Bradley-Terry style pairwise classification loss gives you order and nothing else. The elegant scoring rule argument that makes calibration optimal is built on a quantity the deployed reward model does not actually estimate. He does not soften this, and he is describing the system that had shipped to a hundred million people.
Labeler errors, and the information that is not in the room
There are definitely a lot of labeler errors. There's no way you can have humans label these things and have correct rankings all the time, because sometimes there's just not enough information available to the person doing the labeling. The question might involve some code base that the user has on their computer and the labeler has no way to access. John Schulman, on labeler error, 31:40
They let labelers skip questions they cannot answer, "but I think there's still probably a lot of errors, and it's impossible to read a long answer and catch every single mistake."
Three ceilings, then, stacked: the model must guess, the reward model cannot price how badly it guessed, and the humans training the reward model cannot reliably tell. Each one is a separate wall, and only the first is a law of nature.
Retrieval and citing sources (0:32:21)
Part two. He defines the term first: "retrieval in general in the language model context means your language model is accessing some external source of knowledge. Usually you have some set of documents and you're pulling some text into context to respond to a question."
Then the reasons you would want it, and the order he puts them in is the interesting part.
- Current events. "What's going on in the world."
- Information that was never in pretraining. "Not just because it's new, but because it's some private information, something on your computer, or your code base, or your past conversations."
- Verifiability, which he singles out: "actually I would say the thing that I find even more important, the most important reason for retrieval and citing sources, is verifiability."
The proof sketch analogy
The verifiability argument is not about the answer being more likely to be right. It is about somebody being able to check.
A human has to check responses that models are writing and decide if they're correct or not, and it's extremely hard to check if something is correct if you don't know where the information came from. You have to look everything up. If the model cited its sources it's much easier to check. John Schulman, on verifiability, 33:10
You can think of an unsourced answer as almost like a sketch of a proof. It's kind of like a claim that I have sources to back up all these things but I'm not going to show them to you. John Schulman, on unsourced answers, 33:40
And then the move that connects part two back to part one, which is the thing most people miss about this talk. The citations are not primarily for the user.
Even if we're not going to show sources at test time, when we deploy a model, it's extremely useful in training to be able to get sources so a human can check the information. It's like seeing a full proof instead of a proof sketch. John Schulman, on citations as a training tool, 34:10
Follow the loop. Part one ended with three ceilings, the worst of which was that labelers cannot reliably grade long technical answers. Citations lower that wall. A human grading a cited answer is doing verification rather than recall, which is a different and much easier job. Better labels mean a better reward model, which means a better policy. Retrieval is not a patch applied downstream of RLHF, it is an upstream intervention on the quality of the training signal.
WebGPT: ELI5, a DSL, and 4,000 tokens
WebGPT was "a project that predated ChatGPT," aimed at a narrower kind of question answering.
The data came from ELI5, a dataset built from the Explain Like I'm Five subreddit, and his explanation of why that subreddit is the right source is a nice piece of problem selection:
People ask questions they're curious about, usually it's something that's a little too hard to just Google. If you ask questions that have short clear cut answers, Google will give you a really nice answer box that's probably from Wikipedia that answers your question. For things that are a little more complicated, ELI5 has that kind of question. John Schulman, on the ELI5 dataset, 34:41
The examples on the slide, which he reads out partly because he cannot read them either: something about being in a Zoom meeting on a MacBook, and "why do people recommend baking soda and vinegar as a cleaning agent, that's an interesting one."
The output he shows is the answer to why was the Suez Canal blocked in March 2021, with a couple of different sources and a citation on every claim it makes.
He dates the work and puts its capability in perspective without defending it: "this project was a year and a half ago, two years ago, so this was a GPT-3 level model. I think if you gave a lot of these questions to GPT-4 or even 3.5, it would just answer the questions perfectly without needing to look anything up. But this kind of thing was much more necessary for GPT-3 level models, and it's still useful for GPT-4 for going to more technical esoteric topics."
The action space. This is the design that everything since has copied, so the detail matters. They defined a whole action space, a domain specific language, that the model uses to browse its sources:
- search: issue a query, and the model sees a list of links with little snippets, like a search results page.
- click: open a link.
- quote: save a passage.
- Browser operations, and then a terminal action: when it is done it says I am done and writes its answer.
And the reason quote exists is a hard constraint, stated in the numbers of the time:
Language models have a limited context window, something like 4,000 tokens, each token is about one word, so if you're going to look at a lot of material you're going to run out of space. So quoting is really important. We're only going to be able to show these pages to the model briefly and then we're going to have to move it out of context, so we allow the model to quote the content and that saves it for the rest of the browsing process. John Schulman, on why quoting exists, 37:16
Quoting is a memory primitive, invented because 4,000 tokens cannot hold a research session. That is the whole origin of the pattern, and it is worth noticing that the constraint it answers has moved by three orders of magnitude since while the primitive has not gone away.
One implementation detail he is careful about: "we just defined an RL environment like that where the model emits text. It's not emitting special actions, but the text defines a DSL." No new action heads, no architectural surgery. The action space is a text convention, enforced by training.
The RL task and the pipeline (0:38:05)
The episode, stated exactly: "the model browses for 20 to 100 steps, it quotes a few things, then it writes an answer, and then the reward is computed with the reward model, and we used some standard methodology for this."
Then, at 38:47, he puts up the pipeline picture he had been deferring all talk: "I haven't talked much about the pipeline for RL from human feedback, but here's a picture of it."
The labeling interface that produced one bit
Then a detour into the unglamorous half of the work, which he treats as a finding rather than an aside. "We have to make these GUIs for each of these things." For collecting demonstrations, one interface. For reward modeling, something much heavier, because "we have to get people to read the model written responses very carefully."
The reward modeling interface had labelers highlight statements with strong and weak support. "We had a pretty complex UI for this. I'm not sure exactly how necessary all this stuff was, but we decided to go overboard on defining a really detailed process that people should go through to compute the factual accuracy of the answer."
At the end of the day, after they go through this process of highlighting everything, we just get a binary, we get one bit of information at the end. And we tried using all the other information and it didn't help very much. So that's one disappointing thing. John Schulman, on the labeling interface, 39:48
A deliberately elaborate annotation protocol, and the structured output was worthless. The process was load bearing only as a way of making the single bit more reliable.
The results: best of 64, and the comparison he does not believe
The plots he shows are best of n rather than reinforcement learning: "for a given query you take n samples, you re-rank them with the reward model and you return the best one. So you're not fine tuning, we use the policy from supervised learning, we don't train it with RL."
The headline result, with the numbers exactly as he gives them:
For the biggest model, this is GPT-3, the classic GPT-3, and the most samples, 64 samples, we could beat the human demonstrators. It was preferred 55 to 40 percent of the time, a little worse on coherence but better on factual accuracy. John Schulman, on the WebGPT results, 40:50
They were also preferred a bit over the reference answers, which were written by redditors. And then he takes it back:
Actually I don't totally believe that comparison. I think sometimes people prefer, like, the model writes things that sound very definitive and have all these nice square bracket citations. Even though we didn't tell our labelers, I think we might have even stripped some of the citations out, but the labelers just really like the style of the answers, and I think that biases the comparison unfairly. So I didn't believe that this was actually better than the top upvoted Reddit answers. John Schulman, disowning his own headline comparison, 41:20
He is describing reward hacking at the level of the evaluation rather than the model: a style that reads as authoritative scores above substance, even when the surface markers of authority have been partly removed. The same bias runs through the reward model, which is trained on the same human preferences. The last word is still forward looking: "I think probably if we ran this again with our current models it would be better."
Browsing in ChatGPT, live that morning (0:42:15)
At the time of the talk there was an alpha browsing product in ChatGPT, "kind of using the exact same actions, same sort of methods" as WebGPT. He demonstrates it with a question he asked that morning: who is presenting at the Berkeley colloquium today?
It answers: today's presenter is John Schulman.
Then he opens the debug window, which is the actually informative part. The model is prompted with a long series of instructions about the browser tool and its functions, search, quote, back, "and it describes the documentation for each of the functions."
And in the conversation itself, the model emits an inner monologue alongside each action:
- The user message: who is presenting at the colloquium.
- The assistant's monologue: "I will search for presenter at the Berkeley colloquium today." His own assessment: "it's not very useful, but it tells you what it's thinking."
- The search command, with the recency filter:
Berkeley EECS colloquium presenter today, recency_days=1. And an implementation note dropped in passing: "we use Python syntax now." - The next monologue: "let's click on the first link to access the department colloquium series page for EECS UC Berkeley."
- The click action.
- Then it quotes the relevant passages, and finally writes its answer.
Browsing only when it does not know (0:43:55)
He then names the one thing he thinks is distinctive, and it closes the loop back to part one of the talk:
There are other products that do browsing now and have similar citations. The one thing I think is special about this is that it actually doesn't always do browsing, it only browses when it doesn't know the answer. And I think that uses the same kind of self knowledge of uncertainty that I was describing earlier. The same thing that allows the model to say I don't know allows it to realize it should only browse when it needs to. John Schulman, on selective browsing, 43:56
The demonstration of that is a pair of questions chosen with care, and the second one is a joke aimed at the room.
First: what is the DAgger algorithm? DAgger, dataset aggregation, is "this kind of classic algorithm for imitation learning," and it is also exactly the sort of thing this audience knows. The model "gives a detailed answer, it doesn't browse at all."
Second: he went to the BAIR blog and took the top post, which was about something called Fleet-DAgger. So he asks what is Fleet-DAgger? Now the model does not know, "so it goes and does a search, then it looks at the web page which is actually the full arXiv paper, and then it writes a summary of what Fleet-DAgger is, which I verified is actually a summary. It's not like, it didn't just copy and paste the whole thing, it rephrased it a little bit."
Worth noting what he picked. The top post on the BAIR blog that month was Interactive Fleet Learning, published 6 April 2023, thirteen days before the talk, covering Fleet-DAgger: Interactive Robot Fleet Learning with Scalable Human Supervision. Its authors are Ryan Hoque, Lawrence Yunliang Chen, Satvik Sharma, Karthik Dharmarajan, Brijen Thananjeyan, Pieter Abbeel and Ken Goldberg, which is to say: the man who introduced him, and the professor who convened the lecture series he was speaking in. The demo is a local in joke and a perfect test case at the same time, because it is a real concept that is unambiguously outside a September 2021 cutoff.
"Okay, so that's all for that part of the talk."
Open problems (0:45:35)
I'm at six o'clock now so I'm going to wrap up pretty soon. I wanted to talk a little bit about open problems that I see in this whole line of work. John Schulman, opening the last section, 45:28
Fifteen minutes, three problems, in increasing order of how speculative he is willing to be.
One: getting calibrated uncertainty into words
"One big open problem is just how to incentivize the model to really accurately express its uncertainty in words, and that means using the right amount of hedging, and just explaining its full state of knowledge as well as possible."
He restates why the current method cannot do it, and this time he gives the loss function:
We train the reward models with maximum likelihood, where the probability that A wins over B, our model is that the probability that A wins is proportional to the exponential of the reward score difference. So it's just a kind of classification loss, and it doesn't penalize the model for making extra confident errors, it doesn't account for hedging. I think there's probably some effect where an unhedged wrong answer will be judged as worse than a hedged one, but I don't think we're scoring things exactly right. John Schulman, on the reward model loss, 46:01
That is the Bradley-Terry model, named here in equation form rather than by name. The hedging penalty is not absent, it just arrives indirectly, through whatever the labelers happened to feel, rather than by construction.
And then he argues against the obvious fix, which is the best three minutes of the section. Suppose you did want to train against a real proper scoring rule. You would have to ask the model to output probabilities on everything: 10 percent on this sentence, 20 percent on that one. The problem is not engineering.
That would also have some problems, because natural language is just very imprecise, and that's what makes it powerful. There's just as much fuzziness in the sentence as whatever probability you're assigning. Depending on how you interpret the sentence, there's some underlying interpretation of it, and there are so many possible interpretations, some would have low probabilities, some would have high probability. That makes it very hard to do this. John Schulman, on why you cannot just attach probabilities to sentences, 47:34
A probability is only meaningful relative to a proposition with fixed truth conditions, and a natural language sentence does not have those. It has a cloud of readings with different probabilities. Attaching 70 percent to it is a category error dressed as precision. And the vagueness is not a defect to be engineered away, it is the property that makes natural language useful.
Two speculative directions, both offered as guesses:
- "Maybe we should have some kind of formal statements of probability alongside the natural language statements, but I don't know exactly how to do that."
- "Or maybe we should set up some kind of objective where you have multiple agents collaborating, and they should express uncertainty correctly because it's useful to the other agent, or it's useful to itself later, something like that."
Two: scalable oversight (0:48:45)
The second problem is the labeler ceiling, stated as a research direction: "how do we go beyond things that the labelers can easily do. It's just very hard to check a full long answer about a technical subject or some niche subject." He names the field: "there's this general research area called scalable oversight in the alignment community."
The theoretical hook is a complexity theory argument, and he builds it carefully:
It's often easier to verify that a solution is correct than to generate a correct solution. This is like one of the most basic ideas in theoretical computer science. If you look at the P versus NP problem, one interpretation is that you can have a weak agent, your verifier, that provides an incentive to the strong agent, so that when you optimize the strong agent you're solving a hard class of problems. Say SAT, SAT is the canonical problem that's easy to check the solution but hard to find the assignment. So you can have a weak agent that only does a little bit of compute, that provides the reward, and that'll lead to solving a hard problem if you optimize your strong agent. John Schulman, on verification and generation, 49:07
The conclusion he draws from it: "it seems like it should be possible to have labelers train a model to do things that are much too hard for the labelers to do themselves. In principle it should be possible to do this." Note the double hedge on "in principle."
Two families of approach:
- Decomposition and delegation. "You can try to decompose the task a lot and delegate it, like have your browsing model fact check each sentence and then automatically aggregate all the results." This is part two of the talk turned into a grading instrument.
- Mechanism design. "You can set up some kind of game where you have competing agents that are competing for approval of your verifier, and trying to say why the other is wrong. And there's a nice idea there called AI safety via debate."
His assessment of the state of it, in April 2023:
There's some work in this direction, it's all pretty new, and I think we still have yet to see really good practical implementations of this stuff. But it's starting to become necessary, because it's getting really hard for labelers to keep up with the models. John Schulman, on scalable oversight, 51:08
Three: optimizing for correctness rather than approval (0:51:50)
The third one he flags as "most speculatively," and it is the deepest objection to the method he is best known for.
One unsatisfying thing about RL from human feedback is this purely optimizing on human approval. We don't always know the right answer, and we're probably wrong about lots of things. So we're just optimizing for what sounds convincing and what sounds right, what's kind of the knowledge of the day. It would be great if we could optimize for actual truth, and somehow add more compute and train the models harder and have them get closer to the real truth. John Schulman, on the limit of human approval, 51:38
"So how do you do that?" Two sources of ground truth that are not a human's opinion:
- Prediction. "If you have some kind of ground truth you can optimize for actual correctness. So for example predicting the future. There's a million predictions about the future you can make, and if we use that as the reward function we might be able to generate real knowledge and have a real test for the knowledge. So prediction is one source of generating knowledge."
- Deduction. "You can also obviously do deduction. If you have some kind of formal system or semi formal reasoning system you can generate new knowledge by deduction. So getting our models to do that is another interesting challenge."
"All right, that's all, thanks for your attention."
The failure modes, collected
Scattered across the hour are six distinct failure modes, each with a cause he names and, in most cases, a mitigation he proposes. Collected in one place, with his own hedges intact:
| Failure mode | Cause, as he gives it | His mitigation, and how confident he is |
|---|---|---|
| Hallucination from pattern completion | The model does not know it is allowed to express uncertainty or challenge a premise, and once it has erred it continues coherently | "Pretty easy to fix": a few supervised examples that use the vocabulary. Unlocks the behaviour, does not place it. |
| Hallucination from behavior cloning | The correct target depends on what is in the network, which the data collector cannot see | Reinforcement learning on a reward that prices confident errors. Demonstrated on short answers; not rigorously demonstrated on long ones. |
| Withholding what it does know | Labelers write "I do not know" as a target for facts the network has | The same reward, from the other side. He calls the trade off with informativeness "unavoidable". |
| Reward model cannot price magnitude | Pairwise classification loss recovers which answer wins, not by how much | Open. A proper scoring rule over sentences fails because sentences have no fixed truth conditions. |
| Labeler error at scale | Long technical answers, missing context such as the user's own code base, and the impossibility of catching every claim | Citations so grading becomes verification; skipping unanswerable items; scalable oversight, "all pretty new". |
| Preference for authoritative style | Labelers favour definitive prose and bracketed citations over substance | Named, not solved. It is why he disowns his own Reddit comparison at 41:20. |
Read down the right hand column and the shape of the talk is clear. One problem is solved cheaply, one is solved in the short answer case and asserted in the long answer case, and four are open. He says so each time.
Questions from the room
Five questions, and they are worth having in full, because three of them produce material that is not anywhere in the prepared talk.
Creativity, and combining two patents (0:53:15)
Q: It seems to have an element of creativity. You give it a patent or a pair of patents and say put these together and come up with something new, a new invention, and it seems to do reasonably well with that. Does that surprise you, or how do you consider that new knowledge?
A: "Yeah, that seems like it could be new knowledge. I guess there's some taste that you'd be injecting by asking it that question in the first place, either that it's a good idea to combine inventions, or that these are particularly promising inventions to combine. So it's like you're collaborating with the model to create knowledge to some extent."
And then the line, which lands differently depending on which side of the argument you came in on:
There's not a fine line between creativity and just kind of learning, pattern recognition, pattern completion. John Schulman, on creativity, 55:01
He is deflating the question in one direction and the skeptic in the other. The novelty is real, some of it is contributed by the person choosing the prompt, and the distinction between generating and recombining is not sharp enough to carry an argument.
What is beauty, and whether the model should have opinions (0:55:10)
Q: From someone working with the models on classical literature and philosophy. On a question like what is beauty, where there is no obvious fixed answer but there are many answers, how do you evaluate whether these quantitative measurements of the relative merits of different answers have any standing?
A: He starts by folding it into the problem he has already described. "I talked about the difficulty of judging answers even when they're supposed to be objective and they're not values loaded. So if you have something that's going to depend on taste and values, then that's much harder. I don't think we have a good answer for that."
Then the policy, which is a direct statement of design intent and reads as a period document:
The direction we've been going so far is, we don't think the model should have opinions on things yet, so we want the model to instead be able to describe the set of opinions that humans have. I would want the model to sort of redirect that into a more factual question about what are some human theories, what are the schools of thought that humans have on this. John Schulman, on whether a model should have opinions, 56:01
Note the "yet." The position is explicitly provisional, and it is a stance about the pipeline as much as about ethics: a question with no reference answer has no reward signal, so the model is pointed at the adjacent question that does.
Inner monologue as interpretability, and shorter horizon feedback (0:56:45)
This question comes with a preamble from the audience that is the warmest moment in the hour:
I just want to give you props. I think it might have been five or six years ago I participated in an AI progress forecasting meeting with John, and he was the only person in the room more bullish than me on predicting AI progress. I think he deserves a lot of credit, not only for building what he's built, but for having optimism years in advance that this kind of thing was possible. An audience member, before asking the question, 57:06
Q: On the WebGPT demo, it is really great how it gives this inner monologue. What is your level of optimism versus skepticism for using that inner monologue format for interpretability? Can you distill a model so that it does not have enough room in its inner layers to think and it needs the inner monologue, so we might be able to read out its thoughts?
A: Enthusiastic, with the caveats named in order. "To the extent that we can't find perfect solutions for interpretability, or for making sure our models are safe or well intentioned, I think this is a really good partial solution and we should do as much of it as possible. It's very helpful for interpretability."
Then the two failure modes, including the one the questioner handed him: "Obviously you can't completely trust it. The model could be producing a deceptive inner monologue, that's definitely a concern. But like you said, you could also use a small model so it has to use the inner monologue to reach a certain level of intelligence. Of course then you could worry that it's doing some kind of steganography and it's hiding information, but that's a little far fetched. So overall I think it's promising, but maybe there are some theoretical concerns with it."
And then the part he had not planned to say, which is the best answer in the Q&A. The argument for inner monologue is not primarily interpretability at all. It is credit assignment:
One thing I didn't mention is that if you have detailed inner monologues, that allows you to use a shorter horizon feedback. So for example for browsing, if you don't have the inner monologue and you see one action like scroll, you have no idea if this action makes sense or not, so it's impossible to provide a reward on it. But if the model says I'm scrolling to look for blah and then it has the scroll action, a human can look at that single action and decide if it makes sense or not. John Schulman, on inner monologue as a reward signal, 59:08
By having inner monologue you can train with RL at a shorter time horizon, and that also makes the system safer, because you're not optimizing for long term behavior which could lead to weird results. John Schulman, on short horizon feedback, 59:39
A bare action is unjudgeable. An action plus its stated intention is judgeable, by a human, immediately, without waiting for the episode to finish. The inner monologue makes a long horizon task into a sequence of short horizon ones, which is both easier to train and, on his account, safer.
DAgger versus Fleet-DAgger, and depth of ingrained knowledge (0:59:39)
Q: You mentioned there is some intrinsic knowledge graph in the models, and then you showed the model explaining DAgger versus Fleet-DAgger. DAgger it can explain directly, presumably because the knowledge is inside the model. It is still able to search the web for Fleet-DAgger and explain it, but that has new concepts not inside the model's knowledge graph. Do you expect any difference in the model's capability on the two concepts?
A: Yes, and he labels the epistemic status of his own answer twice while giving it.
I'd say the model is probably best with concepts that it has deeply ingrained, that it's seen in a million contexts. And if it's just seeing the concept for the first time in some document that it's conditioning on, it's probably gonna have less intelligent things to say about it. This is just me kind of half answering based on introspection, or I'm just kind of speculating here. John Schulman, on retrieved versus ingrained knowledge, 1:01:10
I would say it would be better at talking about DAgger than Fleet-DAgger. For Fleet-DAgger it's just going to say some kind of summary of what's in the document, and it's not going to say anything too insightful about it. John Schulman, on the limits of retrieval, 1:01:41
This is a real and underrated limitation of retrieval, and it is the asymmetry the first half of the talk implies: a fact pulled into context is available to be restated, but it has not been integrated into the weights, so the model cannot reason from it the way it reasons from something it has seen a million times. Retrieval buys you accuracy. It does not buy you depth.
The last question: informativeness against correctness (1:01:41)
Q: You mentioned in part one the problem of the model learning to withhold information when that is not desirable. Do you foresee a conflict between the incentive to train the model not to withhold information in open domain contexts, while also training it not to produce unsupported information in closed domain contexts, even when it actually knows that information?
A: The answer is one word longer than it needs to be, and he does not try to resolve it.
Yeah, I think there's an extremely strong conflict. There's a precision recall kind of conflict, there's a conflict between informativeness and correctness, and you often run into this when you're training. With RLHF we're choosing some particular point that we think is reasonable on this trade off curve of how often the model should guess. But it's unavoidable that there's a trade off there. John Schulman, on the last question of the talk, 1:02:11
That is the right place for the talk to end. The first half argued that calibration is learnable. The last answer says that where you put the threshold is a choice somebody makes, it is made once, globally, for every Customer and every context, and the tension it resolves does not go away.
Key takeaways
- Hallucination is not a bug in behavior cloning, it is what behavior cloning teaches. Train on a hundred true answers and the model learns to emit confident answers of that shape, including for the facts it does not hold. The dataset can be perfect and the lesson is still "guess."
- The reason is an information problem, not a quality problem. The correct supervised target depends on what is already in the network, which the person writing the target cannot see. More data and better labelers do not touch it.
- The opposite error is just as real. Write "I do not know" into the targets and you train the model to withhold things it knows.
- A truthfulness dataset is model specific. Targets calibrated to model A teach model B to guess, because the confidences on the edges differ. That is the basis of his April 2023 prediction about models distilled from ChatGPT outputs.
- Pretraining already produces calibrated uncertainty. Log loss is a proper scoring rule, so next token distributions are calibrated, at least for short answers. The uncertainty is in there; the question is only whether the training signal asks for it.
- Under the right reward, hedging below a confidence threshold is simply optimal. With equal reward and penalty the threshold is 50 percent probability on the top choice. His colleague's unpublished TriviaQA experiment recovered exactly that behaviour, and the policy never saw its own log probs.
- The reward model does not need to know the answer, only how well known the answer is. It shares the policy's knowledge graph, which is why it can score hedging at all. Using it in place of the oracle worked, and worked worse.
- The deployed reward model is scoring the wrong thing, and he says so. A pairwise classification loss recovers which answer wins, not by how much, so it cannot price a confident error against a hedged one. That is the gap under the elegant part of the argument.
- Citations are a training instrument before they are a user feature. The point is that a human grading a cited answer is verifying rather than recalling, which raises the quality of the whole signal. He would want sources in training even if they never shipped.
- Retrieval buys accuracy, not depth. A concept arriving in context for the first time gets summarized; a concept seen a million times gets reasoned about. His own hedge: he is speculating.
- Inner monologue is a credit assignment tool as much as an interpretability one. A bare
scrollaction cannot be rewarded.I am scrolling to look for Xfollowed byscrollcan be, which shortens the horizon RL has to optimize over. - Optimizing on human approval optimizes for what sounds right. His proposed exits are ground truth you cannot argue with: prediction about the future, and deduction in a formal system.
- The informativeness and correctness trade off does not get solved, it gets chosen. RLHF picks a point on that curve, once, for everyone. His last words on it: "it's unavoidable that there's a trade off there."
Chapters
The twenty three entries in bold are the video's own chapter markers, reproduced verbatim. The rest are sub beats added here from the transcript clock, because some of the real markers sit minutes away from the material they name, and a few of the gaps run five or six minutes.
- 0:00:00 Introduction
- 0:01:31 Abbeel on recruiting him for Jose Carmena's lab, then keeping him
- 0:02:03 Schulman takes the floor: robotics, then RL, then language models
- 0:02:06 Overview
- 0:02:33 Truthfulness as the biggest technical problem around language models
- 0:03:35 Hallucination, and a note that every example is a first sample, not cherry picked
- 0:04:04 GPT-3.5 Instruct invents the arrest: tigers, and a serval
- 0:04:35 ChatGPT declines the false premise and asks for more information
- 0:05:05 GPT-4 declines, volunteers a September 2021 cutoff, then answers the real question
- 0:05:36 Class one, pattern completion: not knowing it is allowed to be uncertain
- 0:06:08 Reluctance to challenge a premise, getting caught in a lie, then class two, guessing wrong
- 0:06:38 Asking the models about himself, with a contamination caveat first
- 0:07:39 GPT-3.5 sends him to Stanford; GPT-4 adds a math major and misses a year
- 0:08:09 Whether that is even bad depends on where the bio is going
- 0:08:47 Conceptual Model
- 0:09:09 Star Wars triples, and good old fashioned AI
- 0:09:39 A knowledge graph in the weights, with a confidence on every edge
- 0:10:10 Fine tuning as a four line Python function over that graph
- 0:10:42 Why fine tuning is needed at all: the prefix does not say what kind of site this is
- 0:11:24 Behavior Cloning
- 0:11:42 A hundred correct answers still teach the model to hallucinate
- 0:12:12 The Solo example: training it to guess on that type of question
- 0:12:42 The symmetric error, teaching it to withhold what it knows
- 0:13:12 The central claim: the correct target depends on what is in the network
- 0:13:43 The workaround: have labelers ask the model first, then grade the agreement
- 0:14:44 Why it does not transfer, and the prediction about distilled open models
- 0:15:37 Does the model know
- 0:16:16 A precise definition of latent knowledge: a simple piece of code that gets it out
- 0:16:47 Log loss is a proper scoring rule, so pretraining produces a calibrated model
- 0:17:49 The incredulity argument for introspective access, and two cited papers
- 0:18:31 Uncertainty
- 0:18:49 The easy half: a few examples unlock the vocabulary of hedging
- 0:19:41 When should you hedge
- 0:19:53 The reward ladder, five rungs from confident and right to confident and wrong
- 0:20:23 Why it is a proper scoring rule, and why it needs an oracle nobody has
- 0:20:53 The unpublished TriviaQA experiment
- 0:21:55 Behavior cloning first: the model answers one hundred percent of the time
- 0:22:25 What that fine tuning actually taught, which is the format
- 0:22:56 The analytic answer: threshold at 50 percent probability on the top choice
- 0:23:26 RL recovers the thresholding without ever seeing its own log probs
- 0:23:57 Swapping the oracle for a reward model that shares the same knowledge graph
- 0:24:58 It basically worked, and it was worse than the oracle
- 0:25:21 Long form answers
- 0:26:00 Everything is in the grey area; facts are neither right nor wrong
- 0:26:32 Asking InstructGPT about InstructGPT's own reward model objective
- 0:27:32 What a labeler is supposed to do with a half wrong answer
- 0:28:03 The coding case, where he would rather the model guessed
- 0:28:54 Improving factuality
- 0:29:05 The eval: a human checked reference answer, with GPT-4 as the consistency judge
- 0:29:36 Reading the chart: successive ChatGPT versions, then GPT-4
- 0:30:15 Challenges
- 0:30:38 Guessing is unavoidable, and so is misjudging the hedge
- 0:31:09 The ranking loss recovers order, not magnitude
- 0:31:40 Labeler errors, and information the labeler cannot reach
- 0:32:21 Retrieval Citing Sources
- 0:32:40 Current events, private context, and then the reason he cares most
- 0:33:40 An unsourced answer as a sketch of a proof
- 0:34:10 Citations matter in training even if you never show them at test time
- 0:34:41 WebGPT and ELI5: the questions that are too hard to just Google
- 0:35:42 The Suez Canal answer, with a citation on every claim
- 0:36:12 A GPT-3 level model, and what that means for the demo
- 0:36:46 The action space: search, click, quote, and a DSL made of plain text
- 0:37:16 Why quote exists: 4,000 tokens cannot hold a research session
- 0:38:05 RL Environment
- 0:38:25 RL Task, which is 20 to 100 browse steps, a few quotes, then an answer
- 0:38:45 RL Pipeline
- 0:39:17 Comparisons on two whole trajectories, then RL or best of n re-ranking
- 0:39:48 The elaborate labeling interface that collapsed to one bit of information
- 0:40:50 Best of 64 on classic GPT-3 beats the human demonstrators, 55 to 40
- 0:41:20 Why he does not believe the comparison against upvoted Reddit answers
- 0:42:15 Browsing
- 0:42:24 The debug window: the browser tool documentation inside the prompt
- 0:42:55 The inner monologue, the search command, and "we use Python syntax now"
- 0:43:55 Dagger
- 0:44:27 DAgger answered from memory; Fleet-DAgger triggers a search
- 0:45:35 Open problems
- 0:46:01 The reward model loss written out, and what it does not penalize
- 0:47:34 Why you cannot just attach probabilities to sentences
- 0:48:05 Formal probability statements, or agents that need each other calibrated
- 0:48:45 Scalable oversight
- 0:49:07 Verification is easier than generation: P versus NP, and SAT
- 0:50:07 Training a model to do what the labelers cannot do themselves
- 0:50:38 Decomposition and delegation, then mechanism design and debate
- 0:51:08 All pretty new, and becoming necessary as labelers fall behind
- 0:51:50 Optimization for correctness
- 0:52:08 Prediction about the future as a source of real ground truth
- 0:52:38 Deduction as the other source, and the end of the prepared talk
- 0:53:15 Creativity
- 0:54:30 Combining two patents, and the taste you inject by asking
- 0:55:10 Classical Literature Philosophy
- 0:56:01 Why the model should describe the schools of thought rather than hold an opinion
- 0:56:45 AI Progress Forecasting
- 0:57:06 The forecasting meeting, and who in the room was more bullish
- 0:57:38 Inner monologue for interpretability: deception, distillation, steganography
- 0:59:08 The unplanned answer: inner monologue enables short horizon feedback
- 0:59:39 The knowledge graph question, DAgger against Fleet-DAgger
- 1:01:10 Deeply ingrained concepts against one document in context
- 1:01:41 The last question: informativeness against correctness
- 1:02:11 An extremely strong conflict, and a point chosen on the trade off curve
- 1:03:15 Close
Notable quotes
Even if you clone on 100 correct answers, you're teaching the model to hallucinate, because it doesn't have all of those facts. John Schulman, 11:42
The problem with behavior cloning or supervised learning is that the correct target has to actually depend on what knowledge is in the network, and that's unknown to whoever is collecting the data or whoever's doing the experiment. So unless you have a way of looking at what's in the model, you can't train a model to be truthful with behavior cloning. John Schulman, the central claim of the talk, 13:12
There's also the opposite problem, which is that if you try to train the model to say I don't know sometimes, then you're probably gonna also train it to withhold information that it actually has. John Schulman, 12:42
I think if you look really carefully at the factual accuracy, you'd find that they have some problems and they make things up a lot more than the original. That remains to be seen experimentally, but that's what I would predict. John Schulman, on models fine tuned on ChatGPT outputs, 14:44
It would be extremely surprising if it turned out that the model can output a reasonable distribution on that token but it has no introspective access to the uncertainty. John Schulman, 17:49
You can analytically compute what the correct behavior is. Depending on what the penalty is for wrong answers versus the reward for right answers, the optimal behavior is some kind of thresholding where you answer when you have more than 50 percent probability on your top choice. John Schulman, 22:56
The problem about factuality is really not about guessing things wrong or getting things totally wrong. It's about everything being kind of in the gray area. John Schulman, 25:29
I'd rather just have it guess, because that's at least a starting point, I can run the code and debug it. John Schulman, on a coding answer with one wrong argument, 28:03
I don't think our ranking based reward is actually doing exactly the right thing, and that's part of the problem. John Schulman, on the reward model he shipped, 31:09
You can think of an unsourced answer as almost like a sketch of a proof. It's kind of like a claim that I have sources to back up all these things but I'm not going to show them to you. John Schulman, 33:40
At the end of the day, after they go through this process of highlighting everything, we just get a binary, we get one bit of information at the end. And we tried using all the other information and it didn't help very much. John Schulman, on the labeling interface, 39:48
The labelers just really like the style of the answers, and I think that biases the comparison unfairly. So I didn't believe that this was actually better than the top upvoted Reddit answers. John Schulman, disowning his own result, 41:20
The same thing that allows the model to say I don't know allows it to realize it should only browse when it needs to. John Schulman, on selective browsing, 43:56
Natural language is just very imprecise, and that's what makes it powerful. John Schulman, on why you cannot attach probabilities to sentences, 47:34
We're just optimizing for what sounds convincing and what sounds right, what's kind of the knowledge of the day. It would be great if we could optimize for actual truth. John Schulman, on the limit of human approval, 51:38
There's not a fine line between creativity and just kind of learning, pattern recognition, pattern completion. John Schulman, answering a question about combining patents, 55:01
We don't think the model should have opinions on things yet, so we want the model to instead be able to describe the set of opinions that humans have. John Schulman, on whether a model should hold a view, 56:01
If the model says I'm scrolling to look for blah and then it has the scroll action, a human can look at that single action and decide if it makes sense or not. John Schulman, on inner monologue as a reward signal, 59:08
Yeah, I think there's an extremely strong conflict. There's a precision recall kind of conflict, there's a conflict between informativeness and correctness. John Schulman, the last answer of the talk, 1:02:11
Where this sits in the LLM Learning track
This opens the behavior part of the track, and it is the first video in it that is not about building anything. The two before it, the tokenizer build and reproducing GPT-2, end with a trained base model and a loss curve. This talk starts exactly there and says: that object is not yet usable, and the reason is not capability. The full stack tour at the head of the track covers reinforcement learning from human feedback in a few minutes; this is that section expanded to an hour by the person who led the work, and it is the video that explains why an assistant's personality is a training artifact rather than a design document. The three models answering the same false premise three different ways, at 0:04:04, is the proof in one slide.
Watch it before the interpretability talk, which asks a related question from the opposite direction. Schulman shapes behaviour from outside, by pricing it; Olah goes inside and tries to read what the model is doing. Taken together they are the only two strategies available, and the pairing is sharper than either alone: Schulman's whole argument depends on a quantity inside the network that his training signal cannot observe, which is precisely the thing interpretability exists to get at.
It also sets up the engineering stage. The hacker's guide ends on the mistake everyone makes about their own evaluation, and the evals hour answers it with an aligned LLM judge in CI. The judge in that pipeline is the same instrument Schulman describes at 0:29:05, four years earlier and used for the same reason: nobody can read every long answer.
Resources mentioned
The talk itself
- Reinforcement Learning from Human Feedback: Progress and Challenges, the EECS colloquium listing, 19 April 2023, 5 to 6 pm, Banatao Auditorium, 310 Sutardja Dai Hall
- The Berkeley EECS colloquium page for the talk
- ChatGPT architect and Berkeley alum John Schulman on his journey with AI, Berkeley News coverage of the visit
- The spring 2023 Berkeley lecture series on AI, curated by Ken Goldberg, of which this was the fifth talk
People
- John Schulman, and his Berkeley page
- Pieter Abbeel, who gives the introduction and was his PhD advisor
- Jose Carmena, the professor who recruited him to Berkeley to work on prosthetics
- Ken Goldberg, thanked on stage for convening the series, and a co-author of the Fleet-DAgger paper used in the browsing demo
His own algorithms, named in the introduction
- Trust Region Policy Optimization, written at Berkeley with Sergey Levine, Philipp Moritz, Michael Jordan and Pieter Abbeel
- Proximal Policy Optimization, "the most widely used algorithm today in that space, and used as part of ChatGPT's training"
The work discussed
- WebGPT: Browser-assisted question-answering with human feedback, the retrieval project that predated ChatGPT
- Training language models to follow instructions with human feedback, the InstructGPT paper, whose reward model objective is the subject of the long form hallucination example
- Deep reinforcement learning from human preferences, the origin of the comparison based reward model
- GPT-4 Technical Report, where the internal factual eval by category figure lives, and the GPT-4 announcement he reads the chart from
- OpenAI
Datasets and benchmarks
- TriviaQA, the setting for the unpublished thresholding experiment
- ELI5: Long Form Question Answering, built from the Explain Like I'm Five subreddit, the source of WebGPT's questions
- TruthfulQA, the benchmark for imitative falsehood, which is the measurable form of the behavior cloning argument
Papers he gestures at without naming aloud
- Teaching Models to Express Their Uncertainty in Words, Stephanie Lin, Jacob Hilton, Owain Evans, 2022
- Language Models (Mostly) Know What They Know, Saurav Kadavath and colleagues, Anthropic, 2022
The open problems section
- AI safety via debate, Geoffrey Irving, Paul Christiano, Dario Amodei, the mechanism design idea he names
- The Bradley-Terry model, which is the reward model loss he writes out at 46:01
- Proper scoring rules, the frame for both the calibration argument and the reward ladder
- P versus NP and Boolean satisfiability, the verification against generation argument
- Reinforcement learning from human feedback as a general reference
The browsing demo
- DAgger, Stephane Ross, Geoffrey Gordon and Drew Bagnell, the imitation learning algorithm the model answers from memory
- Fleet-DAgger: Interactive Robot Fleet Learning with Scalable Human Supervision, the concept it has to go and look up
- Interactive Fleet Learning, the BAIR blog post from 6 April 2023 he took it from
Other references in passing
- Knowledge graphs, the good old fashioned AI structure his conceptual model borrows
- Solo: A Star Wars Story, the fact outside the cutoff in the worked example
- The 2021 Suez Canal obstruction, the WebGPT answer he shows
- The serval, "that cute cat thing over there," which GPT-3.5 Instruct had him keeping illegally
- Alpaca and Vicuna, the distilled open models released weeks before the talk that his prediction is about
A note on the captions
The automatic captions mangle technical names and most proper nouns in this recording, so names are corrected throughout this page rather than reproduced as transcribed. The corrections, once, for the record:
- Pieter Abbeel appears as "Peter," "theater" and "Peter abiel." Schulman's own pronunciation is clear; the captions are not.
- Trust region policy optimization is transcribed as "translation policy optimization" in the introduction.
- Michael Jordan is the "Mike" in "which he did at Berkeley together with Mike and me," which matches the TRPO author list. That identification is this page's inference from the paper, not something said aloud.
- Ken Goldberg is the "Ken" thanked for hosting the series. The captions give only the first name, and the opening phrase comes out as "I think Kenneth is the fifth in the series," which is almost certainly "I think this is the fifth in the series." Goldberg curated the spring 2023 Berkeley lecture series this talk belonged to, which is the basis for the identification.
- GPT-3.5, GPT-4, ChatGPT and InstructGPT appear variously as "GPD 3.5," "gvd4," "gbd4," "chat EVT," "chat gbt," "instructor gbt" and "instruct gbd."
- DAgger and Fleet-DAgger come out as "dagger," "Fleet dagger," "Fleet Decker" and "Fleet hacker."
- WebGPT is "web GPT," TriviaQA is "trivia QA," ELI5 is "explained like I'm five," the BAIR blog is the "bare blog," and the Berkeley EECS colloquium becomes "Berkeley X colloquium," "Berkeley exclusive" and "Berkeley cloaking."
- Two fragments are not recoverable and are left out rather than guessed: a clause in Abbeel's opening story, and the middle of the audience member's preamble at 0:57:06 about the forecasting meeting, which the captions render as "he was the only person who's for a math camp."
- One apparent error is not an error. "Optimism versus skeptimism, skepticism" at 0:57:38 is an audience member correcting their own slip, and it is quoted here as spoken.
Where it stands, two and a half years on
The talk is a snapshot of a research programme in April 2023, by someone with no way of knowing which of his open problems would be answered. Reading it back from October 2026, the hit rate is unusual, and the misses are as informative as the hits.
The prediction about distilled open models was confirmed in five weeks, by a paper with his introducer on it. The False Promise of Imitating Proprietary LLMs, submitted 25 May 2023 by Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine and Dawn Song, found that models fine tuned on ChatGPT outputs imitate the style convincingly, earn favourable ratings from human reviewers, and barely close the gap on targeted automatic evaluations of the tasks the imitation data does not cover. Both halves of Schulman's argument are in that result: the behaviour transfers and the knowledge does not, and the human preference signal does not notice, because it is reading style. He had named that second failure at 41:20 about his own WebGPT labelers.
Retrieval with citation became table stakes, and so did the judge. The browsing alpha he demos at 42:15 is now the default mode of every major assistant, built from the same primitives: a text DSL, search, fetch, quote, inner monologue. And the automated consistency check at 29:05, where GPT-4 compares a generated answer against a human checked reference, is the technique the industry spent the following two years formalizing under the name LLM as judge, up to and including it being the thing you put in CI.
Scalable oversight stopped being hypothetical. His two families were decomposition and debate. Debate in particular moved from a 2018 proposal he cites to a measured result: Debating with More Persuasive LLMs Leads to More Truthful Answers showed non expert judges reaching higher accuracy when two stronger models argued opposite sides than when reading an answer alone, which is the weak verifier incentivizing the strong agent, working. Separately, his whole list was formalized three months after the talk in Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback, a 30 author survey whose table of contents reads like an expansion of his last fifteen minutes.
The thing he never named is the thing that got named. The word sycophancy does not appear in this talk. The mechanism does, twice: labelers prefer definitive prose, and a reward model trained on labeler preference inherits that. The field caught up in Towards Understanding Sycophancy in Language Models, which traced the behaviour to exactly that source in human preference data.
His third open problem was answered, and not in the way he proposed. "It would be great if we could optimize for actual truth" is, in hindsight, the single most consequential sentence in the hour, because training against checkable ground truth rather than human approval is what reasoning models do. But his two candidate sources of ground truth were forecasting the future and formal deduction, and neither is what the field used. It used mathematics and code, where the answer is cheap to verify and the verifier is a program, and DeepSeek-R1 is the public demonstration of how far that goes. The direction was right and the examples were wrong, which is a very normal shape for a correct prediction.
And the honest misses. The reward model complaint at 31:09, that a pairwise classification loss recovers order and not magnitude, has been worked on from several directions since without being closed; the quantity the elegant scoring rule argument needs is still not the quantity the shipped reward model estimates. Hallucination has not been solved, which he would not find surprising, since he says at 30:08 that some of it is unavoidable in principle. The informativeness against correctness trade off is still a dial somebody sets, and over refusal and overconfidence are still the two ways to set it wrong. On the question that opens the talk, why chat models assert things they have no evidence for, the mechanism he gives in 2023 remains the best short explanation available, and it is still the explanation you reach for when a model invents a citation.
Schulman himself left OpenAI in August 2024, spent roughly five months at Anthropic working on alignment, and has been co-founder and chief scientist at Thinking Machines Lab since February 2025. His own site is the current record.


