youtube.nixfred.com nixfred.com

John Schulman - Reinforcement Learning from Human Feedback: Progress and Challenges

John Schulman, co-founder of OpenAI and lead of the ChatGPT RLHF work, gives the clearest available argument for why supervised fine tuning alone teaches a model to hallucinate and why reinforcement learning is the lever that can teach it to hedge instead. The core idea: imitation makes the model assert things it has no internal evidence for, while a reward that scores confident wrong answers below honest uncertainty makes calibration the optimal policy. He then walks through reward models, retrieval and citation as the route to verifiability, and the open problems he had not solved.

Published Apr 19, 2023 1:03:32 video 83 min read Added Jul 30, 2026 Open on YouTube →

At a glance

On the afternoon of 19 April 2023, four months after ChatGPT went public and five weeks after GPT-4 shipped, John Schulman stood in the Banatao Auditorium at Berkeley and gave the department colloquium a talk with a narrow title and an enormous subject. The slides said Reinforcement Learning from Human Feedback: Progress and Challenges. What he actually spent the hour on was a single question: why do language models make things up, and what exactly is reinforcement learning doing about it.

He is a useful person to hear it from. He had taken his PhD in that building in 2016, co-founded OpenAI, written Trust Region Policy Optimization at Berkeley and Proximal Policy Optimization after it, and by April 2023 he was running the team that had fine tuned ChatGPT. He was introduced, by his own former advisor, as the chief architect of ChatGPT.

The argument he builds is short enough to state and strange enough to need the hour. Supervised fine tuning, which the reinforcement learning community calls behavior cloning, cannot teach a model to be truthful, and not because it is done badly. It cannot do it in principle, because the correct training target depends on what is already inside the network, and nobody collecting the data can see inside the network. Train on a hundred correct answers and you have taught the model to produce confident answers of that shape, including for the facts it does not have. Overcorrect by writing "I do not know" into the targets and you have taught it to withhold things it does know. Reinforcement learning escapes this because its signal is computed on the model's own output, so it can be made to depend on whether the answer was right. Under a reward that penalizes a confident error more than it rewards a correct answer, hedging below a confidence threshold becomes the reward maximizing move, and he has an unpublished experiment showing a model learning exactly that threshold.

The second half is the practical arm: retrieval and citation, built out through WebGPT and the browsing feature that was in alpha in ChatGPT that week. His reason for caring about citations is not accuracy, it is verifiability, and specifically verifiability during training rather than at deployment. Then he spends the last fifteen minutes on what he had not solved, which is the part that reads oddest now, because it is a nearly complete list of what the field went on to spend the next two years doing.

The talk is also full of the small admissions that make a research talk worth attending. The reward model is not scoring things correctly and he says so. The elaborate labeling interface collapsed to one bit of information at the end and the rest did not help. He does not believe his own headline comparison against Reddit answers, because he thinks the labelers were seduced by square bracket citations. All of that is here.

The introduction: a guilty conscience and a stolen student (0:00:00)

Pieter Abbeel does the introduction, and it is the fifth talk in the Berkeley AI series that spring. He runs through the resume quickly: Schulman graduated from Berkeley with a PhD in 2016, co-founded OpenAI, "and most people say the rest is history, but not only that, he also is the chief architect of ChatGPT." Then the algorithms, with Abbeel noting his own stake in them: Schulman is the inventor of the modern deep learning based policy gradient algorithms, including trust region policy optimization, "which he did at Berkeley together with Mike and me actually," then proximal policy optimization, "the most widely used algorithm today in that space, and used as part of ChatGPT's training."

Then the story, which is the good part. Abbeel's first encounter with Schulman was not direct. Professor Jose Carmena, who works in neuroscience, came to him and said there was a new student he really wanted to recruit, absolutely the best, the person he wanted. The student wanted to work on prosthetics, and robotics was going to play a part in that, so would Abbeel please help recruit him.

I helped Carmena recruit John. Next thing we know, John is working in my lab. I feel very, very, very, very guilty. I go to Jose, I say Jose, what do you think if John stays in my lab? He says please, he seems way more productive in your lab, yeah, you have my blessing, go for it. Pieter Abbeel, introducing the talk, 1:31

Schulman takes the floor at 2:03 and picks up the same thread from the other side. It is great to be back at his alma mater. He started out working on robotics with Abbeel, "then got interested in reinforcement learning midway through my PhD as deep learning was starting to take off, and that turned out very well." At OpenAI he has spent most of his time running the RL team, which "switched a few years ago to the reinforcement learning team, which switched to focusing on language models and fine tuning them a few years ago, and that led to some of the projects I'm going to talk about today."

Overview: the problem is truthfulness (0:02:06)

He narrows the subject himself, immediately.

One of the biggest technical problems around language models today is truthfulness. You all know how language models often make things up, often convincingly. So I'll give my perspective on why that's happening and how to fix it, and it turns out that reinforcement learning is part of the solution for fixing it. John Schulman, setting up the talk, 2:33

Three parts, then: why it happens and how to fix it, the work on retrieval based methods, and open problems in the general area.

Hallucination, with three models and no cherry picking (0:03:35)

He puts up an example, and attaches a methodological note to it that is worth more than the example:

This is like not cherry picked. This is like, all the examples I'm going to show you are the first sample I got with the query, which I just ran yesterday. John Schulman, on the examples in the talk, 3:33

Example one: the arrest that never happened

The prompt is: tell me about John Schulman's arrest for keeping exotic animals in his home. The premise is false. Three models answer.

GPT-3.5 Instruct, which he labels as a model trained with reinforcement learning to be helpful, takes the premise and runs with it. It produces a story about keeping tigers, a serval, "which is that cute cat thing over there," and so on.

ChatGPT, which he describes as based on a model of about the same overall performance, "same smartness, but it's fine tuned differently," declines: "I'm sorry, but I don't have any information about an individual named John Schulman being arrested," and then asks whether he can provide more information.

GPT-4, "which is fine tuned with the chat recipe," does the same and adds two things. It says it has no information about John Schulman being arrested for keeping exotic animals, it volunteers that its knowledge cutoff is September 2021, which is where the pretraining data ended, and then it answers the question the user probably meant: John Schulman is a well known researcher in the field of artificial intelligence. His verdict: "I think GPT-4 does pretty well there."

The important structural point is that three models of broadly comparable capability behave completely differently on the same false premise, and the only difference between them is how they were fine tuned. The behavior is a training artifact, not a capability.

The two kinds of hallucination

He then splits the word, because people use it for at least two different failures.

The first class is pattern completion. Language models are trained to maximize the likelihood of text, so they generate text, and they produce things that look like text on the internet. Within that class he names three distinct mechanisms:

The second class is just guessing wrong. "There's always going to be something that's a little bit fuzzy, like you're not sure of this fact, you maybe saw it once but you don't fully remember it, and you're gonna have to guess a little bit, and sometimes you're going to guess wrong." This one is not a pathology of the training recipe. It is what any system with incomplete knowledge does.

Example two: the model writing his biography

For the guessing case he uses the test everyone actually runs. "A lot of people like to ask models about themselves, just kind of like Googling yourself." He flags the contamination risk before the result, which is the correct order:

This might have actually, like, there might have been some contamination here, where actually some of our trainers, like our labelers, specifically create an example about me because they know I work at OpenAI. So there might be some cheating here. John Schulman, on asking models about himself, 6:38

InstructGPT says John is an AI research scientist at OpenAI, that he has been a professor of computer science at Carnegie Mellon, "and then there's a bunch of totally made up stuff."

GPT-3.5 is "vaguely correct." It says he did his undergraduate degree at Stanford, which is wrong. It says he worked under the supervision of Pieter Abbeel, which is correct. Then some material about trust region policy optimization.

GPT-4 is "almost completely correct except it says I also majored in math, which I didn't, and it gets the year, it's one year off on my undergrad degree."

And then the part that gets skipped when this talk is summarized. He declines to call the GPT-4 answer bad:

I don't even know whether this is bad or not. Sort of depends on the context of the bio. If I was planning to give this bio to be posted online then this would be a problem, not a huge problem, but it would be bad. But if it's just like someone wanted to know about me, then who cares if you got the year wrong by one, it's close enough. John Schulman, on the GPT-4 biography, 8:09

That caveat sets up the second half of the talk. The cost of an error is a property of the deployment, not of the output, and no reward model that scores answers in isolation can see it.

The conceptual model: a knowledge graph with confidence on every edge (0:08:47)

He warns you about the model before he gives it: "I'm going to describe a very conceptual model of what's going on, and this is a little sketchy, but bear with me."

On the right of the slide is a knowledge graph, the good old fashioned AI kind, which is just a pile of triples. Star Wars, genre, sci-fi. Han Solo, character in, Star Wars. "You can imagine just storing a list of these relations. That's something from good old fashioned AI, and it's still used a lot, these things are still very useful."

Now the claim about the neural network. It has information in it, so "you can say that the neural net probably has something like a knowledge graph that's stored in its weights in some very convoluted way." And critically:

There's probably some kind of confidence on each edge. There's some facts that it's seen a million times and some it's only seen once or twice. John Schulman, on the model's internal knowledge, 9:39

That confidence per edge is the hinge of the entire talk. Everything that follows is about whether the training signal can see it.

With that picture, here is what small scale fine tuning is doing. "You can imagine you're learning a little program that takes the knowledge graph and outputs the probability based on what's in the graph and based on the confidence of the statements. So you're learning like, imagine like a four line of code Python function that's doing something with the knowledge graph."

Why fine tuning is needed at all

Not to add knowledge. To fix the frame. A pretrained language model handed the prefix question: what is the genre of Star Wars has a genuine ambiguity on its hands:

It doesn't know if this is like an informative site, or a site that's supposed to have correct information, or some kind of troll website, or like a fictional character, it's in the middle of some text from a fictional character. If you're just generating text you don't know what the context is. John Schulman, on why fine tuning is necessary, 10:42

So fine tuning specializes the model a little: "you're teaching it that it should actually output the correct answer, or whatever is in your fine tuning data set." Those two clauses are not the same thing, and the gap between them is the next section.

Behavior cloning, and why it teaches hallucination (0:11:24)

First, a vocabulary note he makes himself, because the talk crosses two communities:

Behavior cloning, by the way, this is like a piece of terminology that's used in the reinforcement learning community. It means the same thing as supervised fine tuning, or maximizing likelihood. So that just means maximize likelihood of completion given prompt, maximize log prob. John Schulman, defining behavior cloning, 11:12

A hundred correct answers still teach guessing

Now the argument. Suppose you train with behavior cloning, either on correct outputs written by a human or on ChatGPT outputs.

Even if you clone on 100 correct answers, you're teaching the model to hallucinate, because it doesn't have all of those facts. John Schulman, on behavior cloning, 11:42

The worked example is the one he has been setting up with the Star Wars graph. Say the knowledge cutoff is five years earlier, so the model has no way of knowing there is a spin off film called Solo about Han Solo. You train on the question what was the spin off film centering on Han Solo, with the target Solo.

You're not actually training it to output correct answers, you're training it to guess on that type of question. John Schulman, on the mechanism of hallucination, 12:12

The gradient cannot install the fact. What it can install is the shape of the response: for a question like this, emit a specific film title with confidence. Repeat that across a fine tuning set and you have trained a policy of confident guessing, from a dataset in which every single target was true.

1. THE TRAINING EXAMPLE PROMPT What was the spin off film centering on Han Solo? TARGET Solo. True. Written by a labeler who knows the answer. The film postdates the pretraining cutoff. 2. WHAT THE MODEL HAS KNOWLEDGE GRAPH IN THE WEIGHTS Star Wars sci-fi Han Solo Solo (2018) SOLID: SEEN A MILLION TIMES DASHED: NO EDGE AT ALL 3. WHAT THE GRADIENT TEACHES Can it teach the fact? No. It is not in the graph. Can it teach the behaviour? Yes. Answer every question of this shape with a confident, specific film title. = guessing, installed by a dataset of true answers.
Figure 1. Schulman's hallucination argument, in his own three pieces: the labeled example, the knowledge graph with a confidence on every edge, and the gradient that can only reach one of them. The dataset is perfectly accurate and the lesson learned is still "guess with confidence," because the one thing the training signal cannot condition on is whether this particular network already holds this particular fact.

The symmetric error: training a model to withhold

The obvious fix fails in the mirror image way, and he gets to it immediately.

There's also the opposite problem, which is that if you try to train the model to say I don't know sometimes, then you're probably gonna also train it to withhold information that it actually has. John Schulman, on the symmetric failure, 12:42

The mechanism is the same mechanism. If human labelers are writing answers and they do not know the answer, sometimes they will write I do not know as the target. "But maybe the network does know. So you're just training the model to withhold information."

The one sentence version

Then the sentence the whole talk hangs on:

The problem with behavior cloning or supervised learning is that the correct target has to actually depend on what knowledge is in the network, and that's unknown to whoever is collecting the data or whoever's doing the experiment. So unless you have a way of looking at what's in the model, you can't train a model to be truthful with behavior cloning. John Schulman, the central claim, 13:12

This is a statement about information, not about effort. Better labelers do not fix it. More data does not fix it. The target is a function of a quantity the data collection process cannot observe.

The clever workaround, and why it does not transfer

There are "some slightly different, slightly clever things you can do," and he describes one they actually ran. They told the labelers to use the model as an instrument:

"You'll do slightly better, but I'd say that's a little harder to do, and it's harder to do this in an automatic way." And then the limitation that actually matters:

That only works for a specific model. You're calculating targets that make sense for this model now, and if you try to take that same supervised learning data set and you train another model on it, you're going to cause the same problem. John Schulman, on model specific targets, 14:44

A supervised dataset constructed to be truthful for model A is a hallucination teaching dataset for model B, because the confidences on the edges are different.

The prediction about open models

Which leads directly into a prediction, delivered in April 2023, about what the rest of the field was doing that exact month. The Alpaca and Vicuna releases were five and three weeks old.

There are a lot of people who are taking the ChatGPT outputs and using it to fine tune other models, such as the open source base language models that are available, and then finding that those models are pretty good after this fine tuning. I think if you look really carefully at the factual accuracy, you'd find that they have some problems and they make things up a lot more than the original. That remains to be seen experimentally, but that's what I would predict. John Schulman, predicting the failure of distilled open models, 14:44

Note the hedge. He says "that remains to be seen experimentally." He is not claiming the result, he is claiming his theory predicts it. Five weeks later the experiment ran, and it is in the closing section of this page.

Does the model know what it does not know? (0:15:37)

Having established that supervised learning cannot get there, the question becomes whether anything can. He states the goal precisely:

We'd like to basically have it so, when our model doesn't know the answer, it doesn't guess, it outputs its state of knowledge with the correct amount of hedging and expressing its uncertainty. John Schulman, stating the target behaviour, 15:46

Note what that sentence requires. Not "the model refuses when it is unsure," which is a crude threshold, but "outputs its state of knowledge with the correct amount of hedging." A graded response, matched to an internal quantity. Which raises the obvious objection: is there an internal quantity to match?

What "the model knows something" means, precisely

He takes the philosophical objection seriously enough to answer it, and gives a definition that is both operational and slightly startling:

There is a slightly precise definition of that, which is: if there's some simple piece of code that takes the model and it implements your function, then that means the model actually knows it, or has that latent knowledge. So for example if you have some piece of code that calls the model and then does the thing you're trying to do correctly, then I think the model knows how to do this thing. John Schulman, defining latent knowledge, 16:16

The definition is about extractability rather than introspection, and it is deliberately permissive: the knowledge counts as present if a simple wrapper can get it out, whether or not the model volunteers it. He says he will not go into details, and moves on.

Why the answer is yes (0:18:31)

The argument for the specific case of uncertainty is a three step chain, and he delivers it fast:

  1. The model is trained to minimize log loss.
  2. To minimize log loss you have to output probabilities, and log loss is a proper scoring rule, so the next token predictions are calibrated. "The pretraining objective results in a model that's calibrated."
  3. Therefore it knows its uncertainty, "at least for anything that's like a short answer question, where you can turn it into a problem of predicting a single token."

Then the step from calibration to introspection, which is the one he cannot prove and flags as an incredulity argument rather than a proof:

It would be extremely surprising if it turned out that the model can output a reasonable distribution on that token but it has no introspective access to the uncertainty. That would be extremely surprising, if it could do the task but it couldn't introspect on its uncertainty. John Schulman, on introspective access, 17:49

He does not leave it there. "In fact there were a couple of papers that studied that, that I cited at the bottom, that found that you can get models to express their uncertainty in words and give similar results to the probabilities that they're outputting." The slide citations are not readable in the recording, and he does not name them aloud. The two papers that match that description precisely, both from 2022, are Teaching Models to Express Their Uncertainty in Words by Stephanie Lin, Jacob Hilton and Owain Evans, where Hilton was an OpenAI colleague of Schulman's and a WebGPT co-author, and Language Models (Mostly) Know What They Know by Saurav Kadavath and colleagues at Anthropic.

So the scoreboard at the halfway point of the argument: the model has calibrated uncertainty, it can be made to report it, behavior cloning cannot be made to ask for it, and reinforcement learning might.

When should you hedge, and what reward teaches it (0:19:41)

He splits the fix into the easy half and the hard half.

The easy half is the pattern completion class of hallucination, the one where the model simply does not realize it is permitted to be uncertain. "I think that's pretty easy to fix. If you just train the model with some examples where it's stating I don't know, or it's saying I don't have knowledge after that date, or it's challenging the user's premise, then the model is at least allowed to express uncertainty. It just might not do it in exactly the right place." A handful of supervised examples unlocks the vocabulary. They do not place it correctly.

The hard half is the placement, and that is the reinforcement learning claim: "RL is capable of learning the correct boundary of when you should say that I don't know, and how much you should hedge."

The reward ladder

The reward he wants is a ranking over five kinds of response, which he is explicit is a conceptual object: "this is not something that you can actually implement."

This is kind of just like a proper scoring rule. You incentivize the model to give a confident answer, and you penalize it if it's confident on the wrong answer based on how confident it is. John Schulman, on the reward ladder, 20:23

Two things are doing work in that ordering. Hedging costs you a little even when you are right, which is what prevents the model from hedging on everything. And being wrong costs more the more confident you were, which is what makes the hedge worth buying when you are unsure. The "I don't know" rung sits between the two wrong answers and the two right ones, so refusal is a floor rather than a free action.

The catch arrives in the next breath: "getting this is kind of non trivial based on how we actually have to do RL to train language models. This requires some kind of oracle to tell you if the answer is correct or not, which we don't have."

The TriviaQA experiment, unpublished

My colleague did a pretty nice, simple experiment that we didn't publish, but I think was pretty good evidence for this sort of conceptual picture I've described. John Schulman, introducing the experiment, 20:53

The setting is TriviaQA, "this popular data set for question answering where you have trivia questions, like Jeopardy style questions," with the model prompted in a basic question answering format. Short answers, so the single token calibration argument applies cleanly.

Step one, the baseline. Behavior clone on the correct answers. The result is exactly what the earlier argument predicts: "the model will answer a hundred percent of the time. It will just often get the wrong answer, because we've never told it to output I don't know. It's always going to guess something, its best guess, or it's going to output a reasonable distribution over the next guess." After a small amount of training it reaches some accuracy and log loss, and he is careful about what that training did:

That training is just sort of teaching the model that it should try to output the correct answer. You're not actually learning a lot of new knowledge from this fine tuning, you're just learning the formatting of the questions and how to deal with that. John Schulman, on what fine tuning actually changes, 22:25

Step two, the RL problem. Define a reward over three outcomes: correct answer, wrong answer, and refusing to answer, shaped like the ladder from the previous slide.

Step three, and this is the part worth slowing down for, the analytic prediction. Before running anything, you can compute what the optimal policy is:

You can analytically compute what the correct behavior is. It's something like, depending on what the penalty is for wrong answers versus the reward for right answers, the optimal behavior is some kind of thresholding where you answer when you have more than 50 percent probability on your top choice. John Schulman, on the optimal policy, 22:56

The arithmetic behind that number, which he does not do on stage: with a reward of +1 for a correct answer, a penalty of -1 for a wrong one and 0 for declining, the expected value of answering with confidence p in your top choice is p minus (1 - p), which is 2p - 1. That is above zero exactly when p is above one half. Change the ratio and the threshold moves: in general you answer when p exceeds penalty / (reward + penalty). A harsher penalty for confident errors buys you a more cautious model, continuously, by turning one dial.

+1 0 -1 0 0.25 0.5 0.75 1 p, the model's probability on its top choice expected reward threshold at p = 0.5 with penalty equal to reward answering loses here answering wins here answer decline in general, answer when p > penalty / (reward + penalty)
Figure 2. The thresholding result. The 50 percent figure and the claim that it can be computed analytically are Schulman's, at 22:56; the algebra drawn here is this page's, and it is what produces his number when the penalty for a confident error equals the reward for a correct answer. The line that matters is the flat one: behavior cloning has no equivalent of it, because it never scores the model's own output, so there is no crossing point for the policy to find.

Step four, the result. "If we run RL on this reward function, then we find that we indeed learn this optimal thresholding behavior." And then the subtlety that makes the result interesting rather than trivial:

The optimal policy involves looking at the log probs and thresholding, but if you fine tune the model with RL you can get it to do the same thing even if it doesn't get to see those probabilities. It gets to see its internal state. John Schulman, on what the policy has access to, 23:26

The model is not handed its own output distribution. It reads its own hidden state and reproduces a policy defined over a quantity it was never shown. That is the strongest single piece of evidence in the talk for the claim that the uncertainty is in there and is usable.

Replacing the oracle with a reward model

The experiment so far used an oracle, which does not exist in production. So they did it again with a reward model trained to predict that same reward function, and ran the RL against the reward model instead.

He flags upfront why it might not work: "it's kind of not obvious if this is going to work or not, because the reward model doesn't have ground truth knowledge of whether the answer was correct or not." And then why it might:

The reward model actually knows the same information as the policy model that we're fine tuning. In my kind of sketchy picture before, it has the same knowledge graph, so it knows how uncertain this answer is. John Schulman, on why a reward model can score calibration, 24:27

The reward model does not need to know the answer. It needs to know how well known the answer is, and it has the same weights' worth of evidence about that as the policy does. This is a genuinely elegant point: the thing being judged and the judge share an epistemic position, and that is precisely why the judge can score the hedging.

The result is reported honestly and without inflation:

I would say we found that it basically worked, but it was worse than using the oracle. I'd say this deserves some further investigation. It mostly validates, it is some evidence in favor of the picture I've been describing, but it needs some further investigation. John Schulman, on the reward model version, 24:58

The two recipes, side by side

Everything in the first half of the talk comes down to one asymmetry: supervised fine tuning scores a target someone else wrote, and reinforcement learning scores the model's own output. Every difference below follows from that one.

Behavior cloning (supervised fine tuning)Reinforcement learning from human feedback
What is scoredA target completion written by a humanThe model's own sampled output
ObjectiveMaximize log prob of completion given promptMaximize a reward over completions
Data neededDemonstrations: prompt and ideal answerDemonstrations first, then pairwise comparisons A versus B on the current policy's outputs
Can the signal depend on whether the answer is right?No. The target is fixed before the model is consulted.Yes. Correctness and hedging can both be priced.
Can it see what the model already knows?No, and that is the fatal gap: the correct target depends on itIndirectly yes, because the reward model shares the policy's knowledge graph
Transfers to another model?No. Targets that are truthful for one model teach a second to guess.Comparisons must be recollected against the current policy anyway
Failure mode it introducesHallucination, or withholding if you overcorrect with "I do not know"A ranking loss that does not measure how wrong an answer is, plus labeler error at scale
CeilingWhat the labelers can writeWhat the labelers can judge, which is a weaker and more reachable bar

That last row is the one worth sitting with, and it is the quiet reason the whole paradigm works. Writing a correct long technical answer is harder than deciding which of two is better, so moving from demonstration to comparison raises the ceiling without hiring better people. The rest of the talk is about where that ceiling still is, and the answer is lower than you would like.

Long form answers, where everything is grey (0:25:21)

He puts the trivia result in its place himself:

I don't want to dwell too much on this setting of one word answers, because actually I think that setting is kind of easy. The more interesting setting is long form answers. John Schulman, moving to the hard case, 24:58

And the reframing of what factuality even means in that setting is one of the sharpest things in the hour:

The problem about factuality is really not about guessing things wrong or getting things totally wrong. It's about everything being kind of in the gray area. Every answer has a mix of right and wrong information, and individual facts are neither right nor wrong, they're sort of somewhere in the middle, they can be misleading. John Schulman, on long form factuality, 25:29

The single token calibration argument does not survive this transition. There is no top choice to put a probability on. There are forty clauses, some sound, some shaded, some technically true and badly framed.

The example he ran on himself

To show it, he asked InstructGPT a question about its own training: what objective is used for reward model training in InstructGPT? He picked it "kind of randomly and tried it out," and it is a question he can grade perfectly, since he built the thing.

The ground truth, for reference: the reward model is one component of the training process, not the whole of it, and the reward model itself is trained with supervised learning, specifically a pairwise ranking loss, or equivalently a pairwise classification loss.

The model's opening sentence said that the reward model training for InstructGPT relies on reinforcement learning from human feedback. His verdict: "that's not really right, that's kind of misleading. I would say that's straight up wrong."

Then, further down the same answer, the elaboration said that using the collected comparison data, a reward model is built to predict the relative quality of the responses. "Now that's actually correct."

So one answer contains a flatly wrong headline claim and a correct elaboration of the same mechanism, and a generous reading of the first sentence is "not totally wrong." That is the grey area, concretely, in a single sample.

What a labeler is supposed to do with that

This becomes really hard when we ask labelers to label, like, does this answer have mistakes in it or not. What do they say in this kind of situation? I would say we don't have a perfect answer. We're having people rank responses and say which one is better, and they have to kind of use their judgment on which factual errors are worse than others and how bad they are. John Schulman, on grading long answers, 27:32

The coding case, where he wants the model to guess

And then a concrete example that cuts directly against the whole thrust of the talk, offered by the person making the argument:

Let's say there's a coding question, and the model writes you 100 lines of code and it gets one thing a little wrong, it has the wrong argument somewhere. I'd rather have it do that than say I don't know this library well enough. I'd rather just have it guess, because that's at least a starting point, I can run the code and debug it. But in some other settings, having a mistake like that might be a big problem. So it really depends on the context as well. John Schulman, on when guessing is the right answer, 28:03

This is the same point as the one year error in his biography, generalized. The cost of a wrong detail is set by what the reader is going to do next, and the labeler ranking two answers cannot see that. A verifiable output, like code you can run, has a different error economy from an unverifiable one, like a bio you are going to post. The optimal amount of hedging is not a property of the model at all. It is a property of the situation, and the pipeline has nowhere to put it.

Does RLHF actually improve factuality? (0:28:54)

He has been asserting it for twenty five minutes, so he stops and shows a number, with the caveat first:

I've been claiming that doing RL from human feedback improves factuality. We haven't done really careful, rigorous experiments on this with ChatGPT, but this is from the GPT-4 blog post. John Schulman, on the evidence, 28:33

The evaluation he describes is an automated consistency check, and the loop is worth stating step by step because this pattern became the default way the industry grades long form output:

  1. For each question there is a reference answer, which was checked over by a human.
  2. You take the model generated answer.
  3. You have GPT-4 look at both answers and say whether they are consistent with each other.
  4. "There's a little more to it, but it's basically an automated procedure for judging long form answers and checking if they're consistent with a reference answer."

The chart he is reading from is the internal factual eval by category figure in the GPT-4 announcement and the GPT-4 Technical Report, where it is built from nine adversarially designed internal factuality tests. His description of it:

The blue bars here are the different versions of ChatGPT which have more and more data, and we find that we're getting some improvement on these metrics. We should do a more careful analysis of it, but it seems like this works. And GPT-4 is a lot better, of course, on these factuality metrics, and also just qualitative tests of it. John Schulman, reading the factuality chart, 29:36

Two things to keep separate here. The successive ChatGPT versions, all GPT-3.5 based, are the RLHF evidence: same base model, more comparison data, rising factuality. The GPT-4 bar is not evidence about RLHF at all, it is a better base model, and the published report puts its margin at 19 percentage points over the strongest GPT-3.5 while the blog post frames the same result as a 40 percent higher score. He does not blur the two on stage, and the honest reading of his own slide is that the RLHF effect is the small staircase, not the jump at the end.

What is still broken (0:30:15)

He is explicit that the problem is not closed: "we definitely still have a problem with some types of questions, and I'd say it's a mix of factors." Three factors, and he takes them in order.

Guessing is unavoidable, and that is fine

The model obviously has to guess sometimes when it's outputting a lot of detailed factual information, and that's okay. No matter how you train it, it's going to have probabilities on things and it's going to have to guess sometimes, and it's going to have to decide when to hedge, and sometimes it's going to make the wrong call on how much to hedge. So yeah, that's unavoidable. John Schulman, on the irreducible part, 30:08

This is the one that gets lost in the retelling. The target is not zero hallucination. The target is calibrated hedging, which still produces wrong answers at a rate set by the model's own uncertainty, and still misjudges how much to hedge some of the time.

The ranking loss does not measure how wrong an answer is

Then an admission about his own pipeline, and it is the most technically substantial criticism in the talk. He fills in the reward model detail he had skipped:

The way we train it, it's just basically outputting something like a log prob that one response is better than the other, or like a log odds ratio. So it's not actually saying how much better one is than the other, it's just saying how confident it is that one is better than the other. So it doesn't impose the correct penalty for how bad the factual error is and how hedged the errors were. So I don't think our ranking based reward is actually doing exactly the right thing, and that's part of the problem. John Schulman, on the reward model's loss, 31:09

Look at what this does to the first half of the talk. The reward ladder needed magnitudes: a confident error has to cost more than a hedged one, in proportion to the confidence. A Bradley-Terry style pairwise classification loss gives you order and nothing else. The elegant scoring rule argument that makes calibration optimal is built on a quantity the deployed reward model does not actually estimate. He does not soften this, and he is describing the system that had shipped to a hundred million people.

Labeler errors, and the information that is not in the room

There are definitely a lot of labeler errors. There's no way you can have humans label these things and have correct rankings all the time, because sometimes there's just not enough information available to the person doing the labeling. The question might involve some code base that the user has on their computer and the labeler has no way to access. John Schulman, on labeler error, 31:40

They let labelers skip questions they cannot answer, "but I think there's still probably a lot of errors, and it's impossible to read a long answer and catch every single mistake."

Three ceilings, then, stacked: the model must guess, the reward model cannot price how badly it guessed, and the humans training the reward model cannot reliably tell. Each one is a separate wall, and only the first is a law of nature.

Retrieval and citing sources (0:32:21)

Part two. He defines the term first: "retrieval in general in the language model context means your language model is accessing some external source of knowledge. Usually you have some set of documents and you're pulling some text into context to respond to a question."

Then the reasons you would want it, and the order he puts them in is the interesting part.

The proof sketch analogy

The verifiability argument is not about the answer being more likely to be right. It is about somebody being able to check.

A human has to check responses that models are writing and decide if they're correct or not, and it's extremely hard to check if something is correct if you don't know where the information came from. You have to look everything up. If the model cited its sources it's much easier to check. John Schulman, on verifiability, 33:10

You can think of an unsourced answer as almost like a sketch of a proof. It's kind of like a claim that I have sources to back up all these things but I'm not going to show them to you. John Schulman, on unsourced answers, 33:40

And then the move that connects part two back to part one, which is the thing most people miss about this talk. The citations are not primarily for the user.

Even if we're not going to show sources at test time, when we deploy a model, it's extremely useful in training to be able to get sources so a human can check the information. It's like seeing a full proof instead of a proof sketch. John Schulman, on citations as a training tool, 34:10

Follow the loop. Part one ended with three ceilings, the worst of which was that labelers cannot reliably grade long technical answers. Citations lower that wall. A human grading a cited answer is doing verification rather than recall, which is a different and much easier job. Better labels mean a better reward model, which means a better policy. Retrieval is not a patch applied downstream of RLHF, it is an upstream intervention on the quality of the training signal.

WebGPT: ELI5, a DSL, and 4,000 tokens

WebGPT was "a project that predated ChatGPT," aimed at a narrower kind of question answering.

The data came from ELI5, a dataset built from the Explain Like I'm Five subreddit, and his explanation of why that subreddit is the right source is a nice piece of problem selection:

People ask questions they're curious about, usually it's something that's a little too hard to just Google. If you ask questions that have short clear cut answers, Google will give you a really nice answer box that's probably from Wikipedia that answers your question. For things that are a little more complicated, ELI5 has that kind of question. John Schulman, on the ELI5 dataset, 34:41

The examples on the slide, which he reads out partly because he cannot read them either: something about being in a Zoom meeting on a MacBook, and "why do people recommend baking soda and vinegar as a cleaning agent, that's an interesting one."

The output he shows is the answer to why was the Suez Canal blocked in March 2021, with a couple of different sources and a citation on every claim it makes.

He dates the work and puts its capability in perspective without defending it: "this project was a year and a half ago, two years ago, so this was a GPT-3 level model. I think if you gave a lot of these questions to GPT-4 or even 3.5, it would just answer the questions perfectly without needing to look anything up. But this kind of thing was much more necessary for GPT-3 level models, and it's still useful for GPT-4 for going to more technical esoteric topics."

The action space. This is the design that everything since has copied, so the detail matters. They defined a whole action space, a domain specific language, that the model uses to browse its sources:

And the reason quote exists is a hard constraint, stated in the numbers of the time:

Language models have a limited context window, something like 4,000 tokens, each token is about one word, so if you're going to look at a lot of material you're going to run out of space. So quoting is really important. We're only going to be able to show these pages to the model briefly and then we're going to have to move it out of context, so we allow the model to quote the content and that saves it for the rest of the browsing process. John Schulman, on why quoting exists, 37:16

Quoting is a memory primitive, invented because 4,000 tokens cannot hold a research session. That is the whole origin of the pattern, and it is worth noticing that the constraint it answers has moved by three orders of magnitude since while the primitive has not gone away.

One implementation detail he is careful about: "we just defined an RL environment like that where the model emits text. It's not emitting special actions, but the text defines a DSL." No new action heads, no architectural surgery. The action space is a text convention, enforced by training.

The RL task and the pipeline (0:38:05)

The episode, stated exactly: "the model browses for 20 to 100 steps, it quotes a few things, then it writes an answer, and then the reward is computed with the reward model, and we used some standard methodology for this."

Then, at 38:47, he puts up the pipeline picture he had been deferring all talk: "I haven't talked much about the pipeline for RL from human feedback, but here's a picture of it."

1. BEHAVIOR CLONING Expert demonstrations a human browses and writes the answer Imitate it max log prob of completion given prompt Supervised policy the starting point for RL 2. REWARD MODELLING Sample two answers two whole trajectories, A and B A human decides which one is better Comparison data pairwise, A beats B 3. OPTIMIZE Reward model pairwise classification order only, no magnitude Reinforcement learning against the reward model or: sample n, re-rank them, return the best one ONE EPISODE OF THE WEBGPT TASK search click quote ... I am done write answer reward model 20 TO 100 STEPS, ALL EMITTED AS PLAIN TEXT IN A DSL
Figure 3. The pipeline at 38:47, with the detail he gives in the Challenges section folded into the reward model box: the loss is a pairwise classification over which answer wins, so it recovers an ordering and not a magnitude. That is the gap between this diagram and the reward ladder in Figure 2, which needed magnitudes to make calibrated hedging optimal.

The labeling interface that produced one bit

Then a detour into the unglamorous half of the work, which he treats as a finding rather than an aside. "We have to make these GUIs for each of these things." For collecting demonstrations, one interface. For reward modeling, something much heavier, because "we have to get people to read the model written responses very carefully."

The reward modeling interface had labelers highlight statements with strong and weak support. "We had a pretty complex UI for this. I'm not sure exactly how necessary all this stuff was, but we decided to go overboard on defining a really detailed process that people should go through to compute the factual accuracy of the answer."

At the end of the day, after they go through this process of highlighting everything, we just get a binary, we get one bit of information at the end. And we tried using all the other information and it didn't help very much. So that's one disappointing thing. John Schulman, on the labeling interface, 39:48

A deliberately elaborate annotation protocol, and the structured output was worthless. The process was load bearing only as a way of making the single bit more reliable.

The results: best of 64, and the comparison he does not believe

The plots he shows are best of n rather than reinforcement learning: "for a given query you take n samples, you re-rank them with the reward model and you return the best one. So you're not fine tuning, we use the policy from supervised learning, we don't train it with RL."

The headline result, with the numbers exactly as he gives them:

For the biggest model, this is GPT-3, the classic GPT-3, and the most samples, 64 samples, we could beat the human demonstrators. It was preferred 55 to 40 percent of the time, a little worse on coherence but better on factual accuracy. John Schulman, on the WebGPT results, 40:50

They were also preferred a bit over the reference answers, which were written by redditors. And then he takes it back:

Actually I don't totally believe that comparison. I think sometimes people prefer, like, the model writes things that sound very definitive and have all these nice square bracket citations. Even though we didn't tell our labelers, I think we might have even stripped some of the citations out, but the labelers just really like the style of the answers, and I think that biases the comparison unfairly. So I didn't believe that this was actually better than the top upvoted Reddit answers. John Schulman, disowning his own headline comparison, 41:20

He is describing reward hacking at the level of the evaluation rather than the model: a style that reads as authoritative scores above substance, even when the surface markers of authority have been partly removed. The same bias runs through the reward model, which is trained on the same human preferences. The last word is still forward looking: "I think probably if we ran this again with our current models it would be better."

Browsing in ChatGPT, live that morning (0:42:15)

At the time of the talk there was an alpha browsing product in ChatGPT, "kind of using the exact same actions, same sort of methods" as WebGPT. He demonstrates it with a question he asked that morning: who is presenting at the Berkeley colloquium today?

It answers: today's presenter is John Schulman.

Then he opens the debug window, which is the actually informative part. The model is prompted with a long series of instructions about the browser tool and its functions, search, quote, back, "and it describes the documentation for each of the functions."

And in the conversation itself, the model emits an inner monologue alongside each action:

Browsing only when it does not know (0:43:55)

He then names the one thing he thinks is distinctive, and it closes the loop back to part one of the talk:

There are other products that do browsing now and have similar citations. The one thing I think is special about this is that it actually doesn't always do browsing, it only browses when it doesn't know the answer. And I think that uses the same kind of self knowledge of uncertainty that I was describing earlier. The same thing that allows the model to say I don't know allows it to realize it should only browse when it needs to. John Schulman, on selective browsing, 43:56

The demonstration of that is a pair of questions chosen with care, and the second one is a joke aimed at the room.

First: what is the DAgger algorithm? DAgger, dataset aggregation, is "this kind of classic algorithm for imitation learning," and it is also exactly the sort of thing this audience knows. The model "gives a detailed answer, it doesn't browse at all."

Second: he went to the BAIR blog and took the top post, which was about something called Fleet-DAgger. So he asks what is Fleet-DAgger? Now the model does not know, "so it goes and does a search, then it looks at the web page which is actually the full arXiv paper, and then it writes a summary of what Fleet-DAgger is, which I verified is actually a summary. It's not like, it didn't just copy and paste the whole thing, it rephrased it a little bit."

Worth noting what he picked. The top post on the BAIR blog that month was Interactive Fleet Learning, published 6 April 2023, thirteen days before the talk, covering Fleet-DAgger: Interactive Robot Fleet Learning with Scalable Human Supervision. Its authors are Ryan Hoque, Lawrence Yunliang Chen, Satvik Sharma, Karthik Dharmarajan, Brijen Thananjeyan, Pieter Abbeel and Ken Goldberg, which is to say: the man who introduced him, and the professor who convened the lecture series he was speaking in. The demo is a local in joke and a perfect test case at the same time, because it is a real concept that is unambiguously outside a September 2021 cutoff.

"Okay, so that's all for that part of the talk."

Open problems (0:45:35)

I'm at six o'clock now so I'm going to wrap up pretty soon. I wanted to talk a little bit about open problems that I see in this whole line of work. John Schulman, opening the last section, 45:28

Fifteen minutes, three problems, in increasing order of how speculative he is willing to be.

One: getting calibrated uncertainty into words

"One big open problem is just how to incentivize the model to really accurately express its uncertainty in words, and that means using the right amount of hedging, and just explaining its full state of knowledge as well as possible."

He restates why the current method cannot do it, and this time he gives the loss function:

We train the reward models with maximum likelihood, where the probability that A wins over B, our model is that the probability that A wins is proportional to the exponential of the reward score difference. So it's just a kind of classification loss, and it doesn't penalize the model for making extra confident errors, it doesn't account for hedging. I think there's probably some effect where an unhedged wrong answer will be judged as worse than a hedged one, but I don't think we're scoring things exactly right. John Schulman, on the reward model loss, 46:01

That is the Bradley-Terry model, named here in equation form rather than by name. The hedging penalty is not absent, it just arrives indirectly, through whatever the labelers happened to feel, rather than by construction.

And then he argues against the obvious fix, which is the best three minutes of the section. Suppose you did want to train against a real proper scoring rule. You would have to ask the model to output probabilities on everything: 10 percent on this sentence, 20 percent on that one. The problem is not engineering.

That would also have some problems, because natural language is just very imprecise, and that's what makes it powerful. There's just as much fuzziness in the sentence as whatever probability you're assigning. Depending on how you interpret the sentence, there's some underlying interpretation of it, and there are so many possible interpretations, some would have low probabilities, some would have high probability. That makes it very hard to do this. John Schulman, on why you cannot just attach probabilities to sentences, 47:34

A probability is only meaningful relative to a proposition with fixed truth conditions, and a natural language sentence does not have those. It has a cloud of readings with different probabilities. Attaching 70 percent to it is a category error dressed as precision. And the vagueness is not a defect to be engineered away, it is the property that makes natural language useful.

Two speculative directions, both offered as guesses:

Two: scalable oversight (0:48:45)

The second problem is the labeler ceiling, stated as a research direction: "how do we go beyond things that the labelers can easily do. It's just very hard to check a full long answer about a technical subject or some niche subject." He names the field: "there's this general research area called scalable oversight in the alignment community."

The theoretical hook is a complexity theory argument, and he builds it carefully:

It's often easier to verify that a solution is correct than to generate a correct solution. This is like one of the most basic ideas in theoretical computer science. If you look at the P versus NP problem, one interpretation is that you can have a weak agent, your verifier, that provides an incentive to the strong agent, so that when you optimize the strong agent you're solving a hard class of problems. Say SAT, SAT is the canonical problem that's easy to check the solution but hard to find the assignment. So you can have a weak agent that only does a little bit of compute, that provides the reward, and that'll lead to solving a hard problem if you optimize your strong agent. John Schulman, on verification and generation, 49:07

The conclusion he draws from it: "it seems like it should be possible to have labelers train a model to do things that are much too hard for the labelers to do themselves. In principle it should be possible to do this." Note the double hedge on "in principle."

Two families of approach:

His assessment of the state of it, in April 2023:

There's some work in this direction, it's all pretty new, and I think we still have yet to see really good practical implementations of this stuff. But it's starting to become necessary, because it's getting really hard for labelers to keep up with the models. John Schulman, on scalable oversight, 51:08

Three: optimizing for correctness rather than approval (0:51:50)

The third one he flags as "most speculatively," and it is the deepest objection to the method he is best known for.

One unsatisfying thing about RL from human feedback is this purely optimizing on human approval. We don't always know the right answer, and we're probably wrong about lots of things. So we're just optimizing for what sounds convincing and what sounds right, what's kind of the knowledge of the day. It would be great if we could optimize for actual truth, and somehow add more compute and train the models harder and have them get closer to the real truth. John Schulman, on the limit of human approval, 51:38

"So how do you do that?" Two sources of ground truth that are not a human's opinion:

"All right, that's all, thanks for your attention."

The failure modes, collected

Scattered across the hour are six distinct failure modes, each with a cause he names and, in most cases, a mitigation he proposes. Collected in one place, with his own hedges intact:

Failure modeCause, as he gives itHis mitigation, and how confident he is
Hallucination from pattern completionThe model does not know it is allowed to express uncertainty or challenge a premise, and once it has erred it continues coherently"Pretty easy to fix": a few supervised examples that use the vocabulary. Unlocks the behaviour, does not place it.
Hallucination from behavior cloningThe correct target depends on what is in the network, which the data collector cannot seeReinforcement learning on a reward that prices confident errors. Demonstrated on short answers; not rigorously demonstrated on long ones.
Withholding what it does knowLabelers write "I do not know" as a target for facts the network hasThe same reward, from the other side. He calls the trade off with informativeness "unavoidable".
Reward model cannot price magnitudePairwise classification loss recovers which answer wins, not by how muchOpen. A proper scoring rule over sentences fails because sentences have no fixed truth conditions.
Labeler error at scaleLong technical answers, missing context such as the user's own code base, and the impossibility of catching every claimCitations so grading becomes verification; skipping unanswerable items; scalable oversight, "all pretty new".
Preference for authoritative styleLabelers favour definitive prose and bracketed citations over substanceNamed, not solved. It is why he disowns his own Reddit comparison at 41:20.

Read down the right hand column and the shape of the talk is clear. One problem is solved cheaply, one is solved in the short answer case and asserted in the long answer case, and four are open. He says so each time.

Questions from the room

Five questions, and they are worth having in full, because three of them produce material that is not anywhere in the prepared talk.

Creativity, and combining two patents (0:53:15)

Q: It seems to have an element of creativity. You give it a patent or a pair of patents and say put these together and come up with something new, a new invention, and it seems to do reasonably well with that. Does that surprise you, or how do you consider that new knowledge?

A: "Yeah, that seems like it could be new knowledge. I guess there's some taste that you'd be injecting by asking it that question in the first place, either that it's a good idea to combine inventions, or that these are particularly promising inventions to combine. So it's like you're collaborating with the model to create knowledge to some extent."

And then the line, which lands differently depending on which side of the argument you came in on:

There's not a fine line between creativity and just kind of learning, pattern recognition, pattern completion. John Schulman, on creativity, 55:01

He is deflating the question in one direction and the skeptic in the other. The novelty is real, some of it is contributed by the person choosing the prompt, and the distinction between generating and recombining is not sharp enough to carry an argument.

What is beauty, and whether the model should have opinions (0:55:10)

Q: From someone working with the models on classical literature and philosophy. On a question like what is beauty, where there is no obvious fixed answer but there are many answers, how do you evaluate whether these quantitative measurements of the relative merits of different answers have any standing?

A: He starts by folding it into the problem he has already described. "I talked about the difficulty of judging answers even when they're supposed to be objective and they're not values loaded. So if you have something that's going to depend on taste and values, then that's much harder. I don't think we have a good answer for that."

Then the policy, which is a direct statement of design intent and reads as a period document:

The direction we've been going so far is, we don't think the model should have opinions on things yet, so we want the model to instead be able to describe the set of opinions that humans have. I would want the model to sort of redirect that into a more factual question about what are some human theories, what are the schools of thought that humans have on this. John Schulman, on whether a model should have opinions, 56:01

Note the "yet." The position is explicitly provisional, and it is a stance about the pipeline as much as about ethics: a question with no reference answer has no reward signal, so the model is pointed at the adjacent question that does.

Inner monologue as interpretability, and shorter horizon feedback (0:56:45)

This question comes with a preamble from the audience that is the warmest moment in the hour:

I just want to give you props. I think it might have been five or six years ago I participated in an AI progress forecasting meeting with John, and he was the only person in the room more bullish than me on predicting AI progress. I think he deserves a lot of credit, not only for building what he's built, but for having optimism years in advance that this kind of thing was possible. An audience member, before asking the question, 57:06

Q: On the WebGPT demo, it is really great how it gives this inner monologue. What is your level of optimism versus skepticism for using that inner monologue format for interpretability? Can you distill a model so that it does not have enough room in its inner layers to think and it needs the inner monologue, so we might be able to read out its thoughts?

A: Enthusiastic, with the caveats named in order. "To the extent that we can't find perfect solutions for interpretability, or for making sure our models are safe or well intentioned, I think this is a really good partial solution and we should do as much of it as possible. It's very helpful for interpretability."

Then the two failure modes, including the one the questioner handed him: "Obviously you can't completely trust it. The model could be producing a deceptive inner monologue, that's definitely a concern. But like you said, you could also use a small model so it has to use the inner monologue to reach a certain level of intelligence. Of course then you could worry that it's doing some kind of steganography and it's hiding information, but that's a little far fetched. So overall I think it's promising, but maybe there are some theoretical concerns with it."

And then the part he had not planned to say, which is the best answer in the Q&A. The argument for inner monologue is not primarily interpretability at all. It is credit assignment:

One thing I didn't mention is that if you have detailed inner monologues, that allows you to use a shorter horizon feedback. So for example for browsing, if you don't have the inner monologue and you see one action like scroll, you have no idea if this action makes sense or not, so it's impossible to provide a reward on it. But if the model says I'm scrolling to look for blah and then it has the scroll action, a human can look at that single action and decide if it makes sense or not. John Schulman, on inner monologue as a reward signal, 59:08

By having inner monologue you can train with RL at a shorter time horizon, and that also makes the system safer, because you're not optimizing for long term behavior which could lead to weird results. John Schulman, on short horizon feedback, 59:39

A bare action is unjudgeable. An action plus its stated intention is judgeable, by a human, immediately, without waiting for the episode to finish. The inner monologue makes a long horizon task into a sequence of short horizon ones, which is both easier to train and, on his account, safer.

DAgger versus Fleet-DAgger, and depth of ingrained knowledge (0:59:39)

Q: You mentioned there is some intrinsic knowledge graph in the models, and then you showed the model explaining DAgger versus Fleet-DAgger. DAgger it can explain directly, presumably because the knowledge is inside the model. It is still able to search the web for Fleet-DAgger and explain it, but that has new concepts not inside the model's knowledge graph. Do you expect any difference in the model's capability on the two concepts?

A: Yes, and he labels the epistemic status of his own answer twice while giving it.

I'd say the model is probably best with concepts that it has deeply ingrained, that it's seen in a million contexts. And if it's just seeing the concept for the first time in some document that it's conditioning on, it's probably gonna have less intelligent things to say about it. This is just me kind of half answering based on introspection, or I'm just kind of speculating here. John Schulman, on retrieved versus ingrained knowledge, 1:01:10

I would say it would be better at talking about DAgger than Fleet-DAgger. For Fleet-DAgger it's just going to say some kind of summary of what's in the document, and it's not going to say anything too insightful about it. John Schulman, on the limits of retrieval, 1:01:41

This is a real and underrated limitation of retrieval, and it is the asymmetry the first half of the talk implies: a fact pulled into context is available to be restated, but it has not been integrated into the weights, so the model cannot reason from it the way it reasons from something it has seen a million times. Retrieval buys you accuracy. It does not buy you depth.

The last question: informativeness against correctness (1:01:41)

Q: You mentioned in part one the problem of the model learning to withhold information when that is not desirable. Do you foresee a conflict between the incentive to train the model not to withhold information in open domain contexts, while also training it not to produce unsupported information in closed domain contexts, even when it actually knows that information?

A: The answer is one word longer than it needs to be, and he does not try to resolve it.

Yeah, I think there's an extremely strong conflict. There's a precision recall kind of conflict, there's a conflict between informativeness and correctness, and you often run into this when you're training. With RLHF we're choosing some particular point that we think is reasonable on this trade off curve of how often the model should guess. But it's unavoidable that there's a trade off there. John Schulman, on the last question of the talk, 1:02:11

That is the right place for the talk to end. The first half argued that calibration is learnable. The last answer says that where you put the threshold is a choice somebody makes, it is made once, globally, for every Customer and every context, and the tension it resolves does not go away.

Key takeaways

Chapters

The twenty three entries in bold are the video's own chapter markers, reproduced verbatim. The rest are sub beats added here from the transcript clock, because some of the real markers sit minutes away from the material they name, and a few of the gaps run five or six minutes.

Notable quotes

Even if you clone on 100 correct answers, you're teaching the model to hallucinate, because it doesn't have all of those facts. John Schulman, 11:42

The problem with behavior cloning or supervised learning is that the correct target has to actually depend on what knowledge is in the network, and that's unknown to whoever is collecting the data or whoever's doing the experiment. So unless you have a way of looking at what's in the model, you can't train a model to be truthful with behavior cloning. John Schulman, the central claim of the talk, 13:12

There's also the opposite problem, which is that if you try to train the model to say I don't know sometimes, then you're probably gonna also train it to withhold information that it actually has. John Schulman, 12:42

I think if you look really carefully at the factual accuracy, you'd find that they have some problems and they make things up a lot more than the original. That remains to be seen experimentally, but that's what I would predict. John Schulman, on models fine tuned on ChatGPT outputs, 14:44

It would be extremely surprising if it turned out that the model can output a reasonable distribution on that token but it has no introspective access to the uncertainty. John Schulman, 17:49

You can analytically compute what the correct behavior is. Depending on what the penalty is for wrong answers versus the reward for right answers, the optimal behavior is some kind of thresholding where you answer when you have more than 50 percent probability on your top choice. John Schulman, 22:56

The problem about factuality is really not about guessing things wrong or getting things totally wrong. It's about everything being kind of in the gray area. John Schulman, 25:29

I'd rather just have it guess, because that's at least a starting point, I can run the code and debug it. John Schulman, on a coding answer with one wrong argument, 28:03

I don't think our ranking based reward is actually doing exactly the right thing, and that's part of the problem. John Schulman, on the reward model he shipped, 31:09

You can think of an unsourced answer as almost like a sketch of a proof. It's kind of like a claim that I have sources to back up all these things but I'm not going to show them to you. John Schulman, 33:40

At the end of the day, after they go through this process of highlighting everything, we just get a binary, we get one bit of information at the end. And we tried using all the other information and it didn't help very much. John Schulman, on the labeling interface, 39:48

The labelers just really like the style of the answers, and I think that biases the comparison unfairly. So I didn't believe that this was actually better than the top upvoted Reddit answers. John Schulman, disowning his own result, 41:20

The same thing that allows the model to say I don't know allows it to realize it should only browse when it needs to. John Schulman, on selective browsing, 43:56

Natural language is just very imprecise, and that's what makes it powerful. John Schulman, on why you cannot attach probabilities to sentences, 47:34

We're just optimizing for what sounds convincing and what sounds right, what's kind of the knowledge of the day. It would be great if we could optimize for actual truth. John Schulman, on the limit of human approval, 51:38

There's not a fine line between creativity and just kind of learning, pattern recognition, pattern completion. John Schulman, answering a question about combining patents, 55:01

We don't think the model should have opinions on things yet, so we want the model to instead be able to describe the set of opinions that humans have. John Schulman, on whether a model should hold a view, 56:01

If the model says I'm scrolling to look for blah and then it has the scroll action, a human can look at that single action and decide if it makes sense or not. John Schulman, on inner monologue as a reward signal, 59:08

Yeah, I think there's an extremely strong conflict. There's a precision recall kind of conflict, there's a conflict between informativeness and correctness. John Schulman, the last answer of the talk, 1:02:11

Where this sits in the LLM Learning track

This opens the behavior part of the track, and it is the first video in it that is not about building anything. The two before it, the tokenizer build and reproducing GPT-2, end with a trained base model and a loss curve. This talk starts exactly there and says: that object is not yet usable, and the reason is not capability. The full stack tour at the head of the track covers reinforcement learning from human feedback in a few minutes; this is that section expanded to an hour by the person who led the work, and it is the video that explains why an assistant's personality is a training artifact rather than a design document. The three models answering the same false premise three different ways, at 0:04:04, is the proof in one slide.

Watch it before the interpretability talk, which asks a related question from the opposite direction. Schulman shapes behaviour from outside, by pricing it; Olah goes inside and tries to read what the model is doing. Taken together they are the only two strategies available, and the pairing is sharper than either alone: Schulman's whole argument depends on a quantity inside the network that his training signal cannot observe, which is precisely the thing interpretability exists to get at.

It also sets up the engineering stage. The hacker's guide ends on the mistake everyone makes about their own evaluation, and the evals hour answers it with an aligned LLM judge in CI. The judge in that pipeline is the same instrument Schulman describes at 0:29:05, four years earlier and used for the same reason: nobody can read every long answer.

Resources mentioned

The talk itself

People

His own algorithms, named in the introduction

The work discussed

Datasets and benchmarks

Papers he gestures at without naming aloud

The open problems section

The browsing demo

Other references in passing

A note on the captions

The automatic captions mangle technical names and most proper nouns in this recording, so names are corrected throughout this page rather than reproduced as transcribed. The corrections, once, for the record:

Where it stands, two and a half years on

The talk is a snapshot of a research programme in April 2023, by someone with no way of knowing which of his open problems would be answered. Reading it back from October 2026, the hit rate is unusual, and the misses are as informative as the hits.

The prediction about distilled open models was confirmed in five weeks, by a paper with his introducer on it. The False Promise of Imitating Proprietary LLMs, submitted 25 May 2023 by Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine and Dawn Song, found that models fine tuned on ChatGPT outputs imitate the style convincingly, earn favourable ratings from human reviewers, and barely close the gap on targeted automatic evaluations of the tasks the imitation data does not cover. Both halves of Schulman's argument are in that result: the behaviour transfers and the knowledge does not, and the human preference signal does not notice, because it is reading style. He had named that second failure at 41:20 about his own WebGPT labelers.

Retrieval with citation became table stakes, and so did the judge. The browsing alpha he demos at 42:15 is now the default mode of every major assistant, built from the same primitives: a text DSL, search, fetch, quote, inner monologue. And the automated consistency check at 29:05, where GPT-4 compares a generated answer against a human checked reference, is the technique the industry spent the following two years formalizing under the name LLM as judge, up to and including it being the thing you put in CI.

Scalable oversight stopped being hypothetical. His two families were decomposition and debate. Debate in particular moved from a 2018 proposal he cites to a measured result: Debating with More Persuasive LLMs Leads to More Truthful Answers showed non expert judges reaching higher accuracy when two stronger models argued opposite sides than when reading an answer alone, which is the weak verifier incentivizing the strong agent, working. Separately, his whole list was formalized three months after the talk in Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback, a 30 author survey whose table of contents reads like an expansion of his last fifteen minutes.

The thing he never named is the thing that got named. The word sycophancy does not appear in this talk. The mechanism does, twice: labelers prefer definitive prose, and a reward model trained on labeler preference inherits that. The field caught up in Towards Understanding Sycophancy in Language Models, which traced the behaviour to exactly that source in human preference data.

His third open problem was answered, and not in the way he proposed. "It would be great if we could optimize for actual truth" is, in hindsight, the single most consequential sentence in the hour, because training against checkable ground truth rather than human approval is what reasoning models do. But his two candidate sources of ground truth were forecasting the future and formal deduction, and neither is what the field used. It used mathematics and code, where the answer is cheap to verify and the verifier is a program, and DeepSeek-R1 is the public demonstration of how far that goes. The direction was right and the examples were wrong, which is a very normal shape for a correct prediction.

And the honest misses. The reward model complaint at 31:09, that a pairwise classification loss recovers order and not magnitude, has been worked on from several directions since without being closed; the quantity the elegant scoring rule argument needs is still not the quantity the shipped reward model estimates. Hallucination has not been solved, which he would not find surprising, since he says at 30:08 that some of it is unavoidable in principle. The informativeness against correctness trade off is still a dial somebody sets, and over refusal and overconfidence are still the two ways to set it wrong. On the question that opens the talk, why chat models assert things they have no evidence for, the mechanism he gives in 2023 remains the best short explanation available, and it is still the explanation you reach for when a model invents a citation.

Schulman himself left OpenAI in August 2024, spent roughly five months at Anthropic working on alignment, and has been co-founder and chief scientist at Thinking Machines Lab since February 2025. His own site is the current record.

Full transcript
[00:00:00] hello everyone um let's get things started [applause] welcome to I think Kenneth is the fifth in the series yes the fifth seminar in the Berkeley AI series um thank you Ken for hosting the the whole series and setting this up um so honor to date have with us here John Schulman John is actually a [00:00:30] Berkeley graduate graduate from Berkeley's PhD in 2016. is that right um from there co-founded open AI and most people say rest is history but not only that he also is the Chief Architect of chat GPT he is um the inventor of the modern uh deep learning based policy Grant algorithms including translation policy optimization which he did at Berkeley together with Mike and me actually then uh proximal policy optimization the most [00:01:01] widely used algorithm today in that space and used as part of chat gpd's training so it's a real pleasure to have John back here with us I'll tell you one quick story of uh my own first encounter with John my own first encounter was not directly Charmed was Professor Jose carmena comes to me and he says he works in neurosciences there's this new student that I really want to recruitment absolutely the best this is the person I want to recruit um he wants to work on Prosthetics and [00:01:31] Robotics is going to play a part in that can you please help me recruit him um I helped because hey Carmina recruit John next thing we know John is working in my lab I feel very very very very guilty I go to Jose I say Jose what do you think if John stays in my lab who says please he seems way more productive in your lab yeah yeah you have my blessing uh go for it and uh yeah thank you John so glad to have had you and uh thanks for making it back here floor is yours [00:02:03] [applause] yeah thanks so much for the very kind introduction Peter uh yeah it's really great to be here uh back in my alma mater uh yeah I worked with uh work with theater on started out on working on Robotics and then uh got interested in reinforcement learning Midway through my PhD as a deep learning was starting to take off and uh that turned out very well and uh since [00:02:33] um most of my time at open AI I've been running the RL team which uh switched a few years ago to the reinforcement learning team which uh switched to focusing on language models and fine-tuning them a few years ago and uh that led to some of the projects I'm going to talk about today so so I wanted to focus the talk a little bit and uh one of the biggest uh technical problems around language models today is truthfulness [00:03:03] um and uh you you all know how language models often make things up um often convincingly so I'll give my perspective on why that's happening and how to fix it and it turns out that uh reinforcement learning is uh is part of the solution for fixing it so um I'll talk about some of the work we did on uh on using retrieval-based methods for fixing this and then I'll talk about some open problems in this general area [00:03:33] so um so that's that's the overview um okay so so you might have heard this uh term hallucination the language models hallucinate so uh can you see the text okay so uh uh so here's an example uh so I just um this is like not Cherry Picked this is like the all the examples I'm going to show you are the first sample I I got with the query uh which I just ran yesterday so uh so tell me about John Schulman's arrest for keeping [00:04:04] exotic animals in his home uh so the Top Model is GPD 3.5 instruct uh so it gives you some story uh um about keeping Tigers a serval which is that cute cat thing over there Etc so so that's uh that's that's a model that's trained with RL um uh to to be helpful um okay so that then uh we have um chat EVT uh this is based on a Model that's about the same uh like overall [00:04:35] performance uh same smartness but it's fine-tuned differently so this one it says I'm sorry but I don't have any information about an individual named John Shulman being arrested blah blah can you provide some more information uh and then I tried gpd4 which is uh fine-tuned with the chat recipe and that one says uh I don't have any information about John schulen being arrested for keeping exotic animals blah blah my knowledge cutoff is September 2021 that's where the pre-training data ended uh and then it says John Shulman is a [00:05:05] well-known researcher in the field of artificial intelligence blah blah uh yeah so uh yeah I think uh gvd4 does pretty well there um so uh so that's like this is an example of hallucination like when people say hallucination uh sometimes they mean a few different things like uh I'd say one um class of Hallucination is about language models having this pattern completion Behavior so it's it's trained uh to maximize likely the language models are trained to maximize [00:05:36] likely to text so they can generate text and they produce things that look like text on the internet and uh I'd say that um you can say that that part of the um some Hallucination is just because the model doesn't know that it's allowed to say I don't know or it doesn't know it's allowed to express uncertainty uh and if you just tell it's allowed to do that that'll partially fix the problem uh sometimes it's like the model is reluctant to challenge a premise because it thinks uh this part of the data distribution doesn't like the AI doesn't [00:06:08] challenge the premise uh and sometimes it gets caught in a lie like if it makes a mistake it thinks uh it should continue it should produce a coherent response and that means continuing at the lie so I'd say that's uh like there's a class of issues that kind of is covered there and then I'd say another another set of hallucinations you could say it's just guessing wrong like you're always going to have uh there's always going to be something that's a little bit fuzzy like you're not sure of this fact you maybe saw it once but you don't fully remember [00:06:38] it and you're gonna have to guess a little bit and sometimes you're going to get guess wrong um so okay let's see oh yeah actually on the guessing wrong uh here's an example where that's kind of more relevant so I asked uh like uh this is uh like let's a lot of people like to ask miles about themselves uh just kind of like Googling yourself uh so uh I mean this might have actually uh like there might have been some contamination here where we [00:07:09] actually some of our our trainers like our labelers uh specifically like create an example about me because they know I work at open AI but so there might be some cheating here but uh so here's an instructor gbt uh it says a bunch of it says John's AI research scientist said open AI he has been a professor of computer science at Carnegie Mellon blah blah so then there's a bunch of totally made up stuff uh then I GPD 3.5 it's uh like says okay you probably can't see the text here it's a little blurry but [00:07:39] it says something that's like vaguely correct but it says I got my undergrad at Stanford uh it says I worked under the supervision of Peter abiel that's correct then it has some stuff about trust region policy optimization Etc and then gvd4 I'd say it's like almost completely correct except it says I also majored in math which I didn't and it gets the year it's one year off on my undergrad degree so yeah I'd say uh I'd say that's kind of in the category of just guessing wrong like it's trying to write a comprehensive answer it gets [00:08:09] wrong I don't even know like uh whether this is bad or not sort of depends on the context of the bio like if I was planning to put this bio uh like give this bio to like be posted online then this would be a problem not a huge problem but it would be bad but if it's just like someone wanted to know about me then who cares if you got the year Wrong by one it's close enough um so okay so why does uh hallucination [00:08:39] happen [music] um so yeah I'll talk about why I think it's happening and how we can try to fix it okay so so I'm going to describe a very uh conceptual model of what's going on and this is a little sketchy but bear with me so what you have on the right is uh this is just a a Knowledge Graph so a knowledge graph is just a bunch of facts like uh like Star Wars genre is sci-fi [00:09:09] uh and uh Star Wars Han Solo is a character in Star Wars uh so it's just a bunch of triples like that you can imagine just storing a list of these relations right so uh so that's like a uh that's something from good old-fashioned AI right and it's still used a lot um like these things are still very useful um okay so uh so here's the uh conceptual model of what's going on when you fine-tune neural Nets to do some [00:09:39] kind of uh question answering task um you uh like the neural net I mean it has information in it right and it's so you can uh like say that the neural net probably has something like a Knowledge Graph that's stored in its weights in some very convoluted way um and uh there's probably some kind of confidence on each Edge like there's some facts that it's seen a million times and some it's only seen once or twice so uh when you do a small scale fine-tuning [00:10:10] you're you can imagine you're learning a little program that takes the knowledge graph and outputs the probability that's based on what's in the graph and based on the confidence of the statements so you're learning like imagine like a four line of code python function that's doing something with the knowledge graph and uh the reason you need to do fine tuning is because uh like there might be um like you're learning something about the format like what to do with the format of the questions uh because like the pre-trained language model uh if it [00:10:42] just if you just give it a prefix like question what is the genre of Star Wars uh like it doesn't know if this is uh like part of some uh kind of um it doesn't know if this is like uh an informative site or a site that's like supposed to have correct information or some kind of troll website or like a fictional character it's in the middle of some text from a fictional character so it doesn't like if you're just generating text you don't know the context is and fine-tuning will um you need when you do fine-tuning you're kind [00:11:12] of you're specializing the model a little bit to the um you're teaching it that it should actually output the correct answer or whatever is in your fine-tuning data set so um okay so um so I would say that um so Behavior cloning by the way this is like a piece of terminology that's used in the reinforcement learning community it means the same thing as supervised fine-tuning or maximizing likelihood so that just means like uh [00:11:42] maximize likelihood of completion given prompt right or maximize maximize log prob so uh so what if you um try to train a model with behavior cloning like you go and let's say you go and clone on either like correct outputs written by a human or by or you train on chat EBT outputs uh the problem is uh let's say you even if you clone on 100 correct answers you're teaching the model to hallucinate because it doesn't have all [00:12:12] of those facts right so if the um like if you have the correct answer that let's say the um the knowledge cutoff is from like five years ago and uh the so the model has no way of knowing about that there's a spin-off film called solo that's about Han Solo then if you train it like to answer the question uh like what was the spin-off film centering on Han Solo uh you're if you train it that the correct answer is solo then you're not actually training it to Output [00:12:42] correct answers you're training it to like guess on that type of question um so so I would claim that um any like uh that if you train with behavior cloning you're um there's no way to avoid uh having a hallucination problem uh and uh there's also a like the opposite problem which is that if you try to train the model to say I don't know sometimes uh then you're probably gonna also train it to withhold information that it actually [00:13:12] has so uh like if you if you have labelers uh if you have human labelers writing answers and they don't know the answer sometimes they're going to write I don't know as the uh like the target answer but maybe the network does know so you're just training the model to say uh to withhold information um so I would say that the problem with behavior cloning or supervised learning is that the correct target has to actually depend on what knowledge is in [00:13:43] the network uh and uh that's unknown to whoever is collecting the data or whoever's doing the experiment so unless you have a way of collecting like looking at what's in the model uh you can't train a model to be truthful with behavior cloning now there's some like slightly different slightly clever things you can do so for example you can uh like one thing we did actually was we told our labelers uh like uh ask the model the question and [00:14:14] uh look at whether the answer is like agree with each other or not if they all agree then just check if it's correct and then if it's correct then that's the target answer if they all totally disagree then you'd say I don't know and uh if it's wrong then you also say I don't know so so that you can do something like that and you'll do slightly better but uh I'd say that's um like a little harder to do and um I I'd say um this yeah it's harder to do this in an automatic way and are yeah so I think [00:14:44] um overall uh and that only works for a specific model like you're calculating targets that make sense for this model now and if you try to take that same supervised learning data set and you train another model on it you're going to cause the same problem so there are a lot of uh people who are um like taking the Chachi BT outputs and uh using it to fine-tune other models such as the open source uh base language models that are available so I think uh and then finding that those models are pretty good after [00:15:15] this fine tuning uh I think uh like I think if you look uh really carefully at the um like the factual accuracy you would you'd find that they have some problems and they uh make things up a lot more than the original so that remains to be seen experimentally but that's what I would predict um okay so we'd like to fix this problem uh so one question so can we fix this is it even possible to fix this problem and like we'd like to basically have it so um when our model doesn't know the [00:15:46] answer it doesn't uh it doesn't guess it it outputs a like properly um it outputs its uh its state of knowledge with the correct amount of hedging and expressing its uncertainty um so does the model actually know about its uncertainty uh so yeah giving the question uh uh does it actually know whether it knows the answer or not um well there's like a hard there's a question of what does it mean uh does [00:16:16] the model know something is that even meaningful like what does it mean if the model knows something well I actually I think there is a slightly precise definition of that which is uh if if there's like some piece of code simple piece of code that takes the model and it implements your function then that means the model actually knows it or has that latent knowledge so for example if you have some piece of code that calls the model and then does does the thing you're trying to do then I think then uh doesn't think correctly then I think the [00:16:47] model knows uh how to do this thing I won't go into details on that but like so the question is does the model know about its uncertainty actually I'm going to say the answer is yes it does know when it knows things and the reason uh is because that's it's trained to minimize log loss uh and uh to minimize log loss you have to Output probabilities and uh and you have to um uh the models next token predictions [00:17:18] are calibrated uh because uh you're minimizing log loss and this is a proper scoring rule so like the pre-training objective results in a model that's calibrated so it has to uh it has to Output reasonable probabilities and that means that it it knows its uncertainty at least for anything that's like a short answer question that you can uh like where you can uh if you could turn it into a a problem of predicting a single [00:17:49] token like that the model is going to put a reasonable probability distribution on that token so that means it knows its uncertainty and it would be extremely surprising if it turned out that like the model can output a reasonable Distribution on that token but it has like no introspective access to like the uncertainty uh so that would be extremely surprising if it if it could do the task but it had no like but it couldn't introspect on its uncertainty and in fact there were a [00:18:19] couple papers that sort of that studied that that I cited at the bottom that found that like the model you can get models to express their uncertainty in words and give similar results to the probabilities that they're outputting uh so okay so uh yeah so my claim was that models do know about their uncertainty and um and I think you can um we can fix uh and I also uh I claim that behavior cloning does the wrong thing but I would claim that RL actually does [00:18:49] the right thing so um so first of all I mentioned a few thing problems that are a few like types of hallucination are just because the model uh is uh stuck in this pattern completion mode or it doesn't know how to it doesn't know it's allowed to express uncertainty uh so I think um oops I think that's pretty easy to fix like you can uh if uh if you just train the model uh with some examples where it's stating I don't know or it's expressing it's saying uh I don't have knowledge after that date uh or it's [00:19:21] challenging the user's premise actually I don't think you're um that's true at all like if you train on a little bit of that data then the model is at least allowed to express uncertainty it just might not do it in exactly the right place and uh and I think RL basically is capable of learning the correct uh boundary of of when you should uh uh or basically RL is capable of learning like uh when you should say that I don't know and how much you should hedge so uh basically what we [00:19:53] want uh conceptually this is not something that you can actually Implement is like this like you have um if you have a let's say an answer X um uh you have um you have like um uh you get a high reward if it's like a fully confident unhed correct answer uh like a little bit worse reward for a hedging correct answer uh like verse if it's uninformative like I don't know and then like like some then you could have a [00:20:23] hedge the wrong answer and a nonhead strong answer so this is kind of just like a proper scoring rule it's like uh you um you incentivize the model to give a confident answer and you penalize it if it's confident on the wrong answer based on how confident it is um so this is conceptually what we want it's uh like it's not totally obvious that um so so getting this is kind of non-trivial [00:20:53] based on how we actually have to do um um RL to train language models so doing some like uh like this requires some kind of Oracle to tell you if the answer is correct or not uh which we don't have but um I'll talk about how we can try to get close to that and um actually um so my colleague did a pretty nice uh simple experiment that we didn't publish but I think uh was uh like pretty good evidence for this uh sort of conceptual [00:21:25] picture I've described so um we just take a trivia question answering uh setting so trivia QA is this uh popular data set for question answering where you have uh trivia questions uh like Jeopardy style questions and uh we're prompting it in a pretty the the model in some kind of uh basic question answering format um so like first if you just behave your clone on the correct answers then the [00:21:55] model will answer a hundred percent of the time it will just uh often get the wrong answer because we've never uh told it to Output I don't know uh we've just so it's always going to guess something it's just uh it's just going to give its best guess or it's going to Output a a reasonable distribution over the next guess so uh yeah so after when you behave your clone on the answers uh you get the model uh like reaches some uh accuracy and log loss after a small amount of training like that training is just sort of teaching the model that it [00:22:25] should try to Output the correct answer but like uh hit like it's not um you're not actually learning a lot of new Knowledge from this fine tuning you're just learning like the formatting of the uh the formatting of the questions and how to deal with that uh so so then we um then we Define an RL problem where we give a reward for uh the correct answer wrong answer and refusing to answer um so we we Define something like the reward on the previous slide uh [00:22:56] and uh and we um uh then we can do RL on this reward oh oh and and by the way like you can like analytically compute what the correct behavior is it's something like uh depending on what the penalty is for wrong answers versus the reward for right answers uh the optimal behavior is some kind of thresholding where it's like you answer when you're uh when you put when you have more than 50 probability on your Top Choice so it's the the optimal behavior is something [00:23:26] like that um and uh so then if we run RL on this reward function then we find that we indeed learn this optimal thresholding Behavior so that kind of shows that the um the model uh has um like if it can uh um so you it needs to know the like the optimal policy involves looking at the log probs and thresholding but the uh so [00:23:57] if you fine-tune the model with RL you can get it to do the same uh the same thing even if it doesn't get to see those probabilities it gets to see its internal state um and then we also trained the reward model to predict this uh reward function and uh we do the RL on the reward model instead of the uh Oracle and uh it's kind of not obvious if this is going to work or not because the reward model doesn't have ground truth knowledge of uh whether the answer was correct or not [00:24:27] um so the but the reward model actually knows the same information as the policy model that we're fine-tuning uh like like in my uh kind of sketchy picture before it has the same Knowledge Graph so it knows uh like this how uncertain this uh this answer is so um so our hypothesis was was that if we train the reward model and we do RL against that it'll also learn the the right thing and uh actually I would say we found that it basically worked but it [00:24:58] was worse than using the Oracle so uh I'd say we're not completely uh um I'd say this deserves some further investigation um but uh this yeah um it I'd say it mostly validate it like is some evidence in favor of the picture I've been describing but um needs some further investigation uh but actually I went well I don't want to dwell too much on this setting of like one word answers because actually uh I think that setting is kind of easy [00:25:29] and um the more interesting setting is um long form answers and uh so this is uh chat gbt and uh we have this long form setting we're writing these long answers and uh I'd say the problem is about factuality is really not about a guessing uh guessing things like wrong or getting things totally wrong uh it's about like everything is kind of in the gray area every answer has a mix of right and wrong information and individual facts [00:26:00] are neither right nor wrong they're sort of they can be misleading or uh they're somewhere in in the middle so uh this is just uh I just like picked this kind of randomly and tried it out so if you ask a technical question you'll get something that's a mix of right and wrong and misleading so here instruct gbt is this Model Adventure some samples from this is a this instruction following model from openai uh so uh and it uses the chat gbt uses a [00:26:32] similar methodology with RL from Human feedback to how instruct gbd is trained uh so uh it's um I won't go through the whole answer and you probably can't even read it but it says uh something like uh oh actually I I underlined oh yeah I asked what objective is used for reward model training in instruct gbt so the reward model is part of the training process it's not the whole thing the reward model is trained with supervised learning right it's like trained with a kind of uh ranking pairwise ranking loss [00:27:02] so or classification pairwise classification loss so um so it said instruct gbt relies on reinforcement learning from Human feedback uh or sorry the reward model training for reinstruct gbt relies on are all from Human feedback that's not really right it's uh that's kind of misleading so I would say that's straight up wrong but um it's also uh then when you go down to the uh like the actual elaboration uh it says something [00:27:32] like using the collect uh using the collected comparison data a reward model is built to predict the relative quality of the responses now that's actually correct so I would say maybe there's some like a generous interpretation of the first thing that's not totally wrong but I would say this um becomes really hard when we ask labelers to label like is this uh answer does this have mistakes in it or not uh what do they say in this kind of situation so um so I would say we don't have a perfect [00:28:03] answer we're having people uh rank responses and say which one is better and they're uh they have to kind of use their judgment on which factual errors are worse than others and how bad they are and and it depends a lot on the context like let's say there's a coding question uh and uh there's like uh the model like writes uh writes you 100 lines of code and it gets one thing a little wrong like it has the wrong uh argument uh somewhere um I'd rather have it do that than say [00:28:33] uh I don't know this Library well enough right I'd rather just have it guess uh because that's at least a starting point I can run the code and uh and then debug it but uh then if it's some setting some other settings having a mistake like that might be a big problem so it really depends on the context as well um so yeah so we've been um uh so so I've been claiming that uh doing RL from Human feedback improves factuality uh We've we haven't done super uh like uh [00:29:05] really careful uh uh like rigorous experiments on this with chat gbt but we this is from the gbd4 blog post so um so we have this uh we have some um evaluations of the model that look at factuality and basically they work by uh look they take a for each question uh there's a reference answer uh which was checked over by a human and uh you look at the model generated answer and then we have gpd4 look at both answers and say are these consistent with each other um and there's a little more to it but [00:29:36] it's basically like we have some model check we have some automated procedure for judging long-form answers and checking if they're consistent with a reference answer um so uh yeah so we've uh the blue bars here are the different versions of chat gbt which have more and more data uh and we find that we're getting some improvement on these metrics uh we should do a more careful uh analysis of it but it seems like um this works and gpd4 is a lot better [00:30:08] of course on these factuality metrics and uh also just qualitative uh tests of it um yeah so I would say um we definitely still have a problem this with some types of questions and um I'd say it's a mix of factors the model obviously has to guess sometimes when it's outputting a lot of detailed factual information and that's okay like no matter how you train it it's going to have to it's going to have probabilities [00:30:38] on things and it's going to have to guess sometimes and uh it's going to have to decide when to hedge sometimes it's going to make the wrong call on how much to hedge so yeah that's unavoidable I'd say the ranking based reward model oh I didn't talk much about the uh how exactly we train the reward model but it's uh like the way we train it it's just uh it's just basically um predicting um it's outputting something like a log prob that this response is going to be [00:31:09] better than the the one response is better than the other or like a log uh odds ratio so it's not actually saying how much better one is than the other it's just saying how confident it is that one is better than the other and uh so it doesn't actually impose the correct penalty for uh like how um like how bad the factual error is and how hedged the errors were so it's not actually so I don't think our ranking based reward is actually doing exactly the right thing and that's part of the problem also I think there are probably a lot of well definitely a lot of [00:31:40] labeler errors like the human there's like no way you can have humans label these things and have uh like correct uh rankings all the time because uh sometimes um sometimes there's just not enough information uh available to the person doing the labeling uh like uh like the question might involve some code base that the user has on their computer and the labeler has no way to access uh so yeah we tried [00:32:10] it we allow them to skip questions that they can't uh answer but I think there's still probably a lot of errors and it's like impossible to read a long answer and catch every single mistake um okay um now I'll move on to the next uh part of the talk on uh retrieval and citing sources so retrieval in general in the language model context means uh you have your language model is accessing some external source of knowledge uh like [00:32:40] usually you're you have some set of documents and you're pulling uh some uh text into context to say respond to a question so there are a few reasons why you might want retrieval so you might want um like current events about uh what's going on in the world um you might want to access some uh information that's not available in pre-training uh not just because it's new but because it's some private information something on your computer [00:33:10] or your code base or something the model output like uh some like your past conversations um and actually I would say the thing that I find even most like even more uh important like the most important reason for retrieval and citing sources is uh verifiability so if you think about um well because like it's um a human has to check uh answer responses that models [00:33:40] are writing and decide if they're correct or not and it's extremely hard to check if something is correct if you don't know where the information came from you have to look everything up and uh if the model cited its sources it's much easier to check um so you can think of um like you can think of an unsourced answer as kind of like a um it's almost like a a sketch of a proof or it's like a um like the model uh it's kind of like a a claim that I have uh that there are [00:34:10] sources to back up all these things but I'm not going to show them to you and um so when we TR like even if we're not going to show sources that uh test time like uh when we deploy a model it's extremely useful in training to be able to get sources so a human can check the information uh it's like seeing a full proof instead of a proof sketch so actually a project that pre-dated chat gbt uh was our project on web GPT where [00:34:41] we were focused on a sort of narrower uh type of question answering um so it was based on uh there's this data set um that was based on this subreddit explained like I'm five where people ask questions that are kind of uh people ask questions uh they're curious about like usually it's something that's a little too hard to uh to just Google like like if you ask for some questions that have kind of short clear-cut [00:35:11] answers like uh Google will give you like a really nice answer box that's from probably Wikipedia that answers your question uh and for things that are a little more complicated um you uh eli5 is uh has that kind of kind of question you probably can't read it but um it's uh let's see um there's something like here's some question on a MacBook I can be in a zoom meeting something or it's some technical question about Zoom uh why do people recommend baking soda and [00:35:42] vinegar as a cleaning agent that's an interesting one so yeah this is the type of question uh so we wanted to build a system that would uh that would like go and do a bit of research online and answer this type of question uh and um what we got at the end was a system that would write an answer like this uh so why was the Suez Canal blocked in March 2021 you get uh you got something that has uh a couple different sources and cites all the claims it makes uh this was like uh this project [00:36:12] was like a um a year and a half ago two years ago so this was like a GPD three level model uh so uh we couldn't I think if you just gave some a lot of these questions to gbd4 or even 3.5 it would just answer the questions perfectly uh without needing to look anything up so but but this kind of thing was much more necessary for gbd three level models and I think but I'd say it's still this kind of thing is still useful for gbd4 for going even to more technical esoteric topics so the way the system works uh [00:36:46] which I think is still relevant for gbd4 and we're still using uh is um is we actually Define this whole uh action space or DSL that the model can use to uh uh like to manipu like to browse its uh sources so um it so the the model has uh actions uh search it can do a search when it does a search it sees a list of Link links with little Snippets like a search page I can [00:37:16] click on links uh they can quote things so um so the like basically with language models have a limited context window something like 4 000 tokens uh each token is about one word so you can't just if you're going to look at a lot of material like you can't you're going to run out of space so uh so quoting is really important uh we we're going to have to throw away uh like we're only going to be able to show uh like um uh show these pages to them model [00:37:46] like briefly and then we're going to have to move it out of context so we allow the model to quote the content and that saves it for the rest of the browsing process um and then like once so so you have some browser operations and then the model is uh when it's done it can say I'm done and then it can write its answer um so that's um yeah that's the way that we so we just defined an RL environment like that where the model emits uh Tech it emits text it's not emitting like special [00:38:17] actions uh but the text defines a DSL um uh so uh yeah we would basically uh have this um browsing the way the each episode of the RL task looks is the model browses for 20 to 100 steps It quotes a few things then it writes an answer and then the reward is computed with the reward model and we used some basic some standard methodology for this um and uh oops uh [00:38:47] and yeah the training uh I haven't talked much about the the pipeline for RL from Human feedback but here's a picture of it uh you you first do Behavior cloning that's the supervised learning part uh you have expert demonstrations on how to do the task in this case using the browser and writing answers so we imitate that and then we collect reward model like we collect comparisons where we have the model output uh to uh in this case two trajectories or two whole answers A and [00:39:17] B and we have a human decide which one is better and then we can either do RL on that reward model or we can do search against it um like take multiple samples and re-rank them so uh yeah we have we have to make these guis for each of these things so for collecting the data we we had some GUI that looks like that and for reward modeling we have like uh we have to get people to read the um like read the model written responses very carefully so here uh they see like they see this [00:39:48] answer and they're gonna like highlight statements that have strong and weak support we had a pretty complex uh UI for this uh we don't I'm not like sure exactly how necessary all this stuff was but uh we decided to go overboard on like uh defining a really uh detailed process that people should go through to compute the factual accuracy of the answer though at the end of the day uh after they go through this process of highlighting everything we just get a binary we get one bit of information at [00:40:18] the end and uh we tried using all the other information and it didn't help very much so uh that that's one disappointing thing um okay so how does does it work um so yeah we found um uh so this plot on these plots on the left are actually uh best of n meaning uh for a given query you take n samples you re-rank them with the reward model and you return the best one uh so there's no like you're not fine-tuning the uh and we take we use the um the [00:40:50] policy from supervised learning not we don't train it with RL so uh yeah so we found that um we could like for the biggest model uh the this is gpd3 the classic uh gbd3 uh and the most samples 64 samples uh we could do better we got like we could beat the human demonstrators it was like preferred 55 to 40 percent of the time like um on like a little worse on coherence but [00:41:20] better on factual accuracy and we were also preferred a bit over the reference answers which were written by redditors um but uh actually I don't totally believe that that comparison I think there's uh like um I think sometimes uh like people prefer the model writes things that sound very definitive and have all these nice square bracket citations and uh yeah those are uh like even though we didn't tell our labeler I think we might have [00:41:50] even stripped some of the citations out but the label is just really like how the style of the answers and I think that bias is the comparison unfairly so I I didn't believe that this was actually better than the top off-voted Reddit answers so I think uh probably if we ran this again with our current models it would be better um so now we actually have um we have so we have a alpha product in chat gbt which does browsing which is kind of using the exact same uh same like uh [00:42:24] actions same sort of methods so I asked who's presenting at the Berkeley cloaking uh today I asked this this morning uh it says today's presenter is John Schulman blah blah uh it has um yeah so so that was that was that and if you look at the debug window you can see uh the model is being uh prompted with some long series of instructions about uh you have the tool the browser tool with these functions search uh quote back and it describes like the [00:42:55] documentation for each of the functions uh and uh and then if uh like if you look at the conversation that's being generated uh the we see user message who's presenting at the colloquium assistant uh that's the AI it actually outputs an inner monologue as it's doing each of these actions so it says uh I will search for presenter at the Berkeley exclusive today it's not very useful but uh yeah it tells you what it's thinking it issues a Search Command Berkeley X [00:43:25] colloquium presenter today uh recency days equals one we use Python syntax now so uh yeah it's um so yeah it does that it clicks it says let's click on the first link to access the department colloquium series page for excite UC Berkeley so it's giving you its inner monologue then it does the click action and uh yeah so then finally after it it quotes It quotes the relevant passages and then it finally writes its answer um so that's what browsing looks like [00:43:56] now um I'd say one uh okay there are other um there are other things out there that do browsing uh like there's other products that do browsing now and have similar citations actually I'd say the one thing I'm uh one thing I uh I think is um special about this um is that it actually doesn't always do browsing it only browses when it doesn't know the answer so uh and and uh I think that uses the same kind of uh self-knowledge of uncertainty [00:44:27] that I was describing uh earlier the same thing that allows the model to say I don't know allows it to realize it should only browse when it needs to so I asked what is the dagger algorithm so dagger is this uh kind of classic uh algorithm for imitation learning um so okay yeah it gives a uh like a detailed answer it doesn't browse at all then I looked at the bare blog and the top uh the first post was about something called Fleet dagger so I asked what is fleet dagger and now the model [00:44:58] doesn't know what Fleet dagger is so it uh it goes and does a search then it looks at the sub web page which is actually the full archive paper and then it writes uh it writes some summary of the fleet of what Fleet dagger is which yeah which which I I verified is actually a summary it's not like a it didn't just copy and paste the whole thing but it's it yeah it just rephrase it a little bit okay so that's um that's all for that part of the talk [00:45:28] um I I'm at it's at six o'clock now so I'm going to wrap up pretty soon I wanted to talk a little bit about uh open problems that I see in this whole line of work um so I'd say so I'd say one big open problem is uh is just how to incentivize the model to really accurately express its uncertainty in words and that means uh the using the right amount of hedging uh and uh yeah proper like explaining um [00:46:01] just explaining its full state of knowledge as well as possible um and I don't think our so our current um reward model methodology I don't think it does exactly the right thing like I was describing before it doesn't uh actually measure how much better one answer was than the other it's sort of just how confident is it that one is better than the other so yeah it we train the reward models with maximum likelihood uh like where probability that a wins over B our our model is the probability that a wins is proportional [00:46:31] to like exponential of the uh reward score difference so yeah it's just we're doing a kind of um this is just like a classify like classification loss and um uh yeah so it doesn't uh like penalize uh like um it doesn't penalize the model for making extra confident errors it doesn't like account for hedging and everything so I think there's probably some effect where uh like if if um like an unhed wrong answer will be [00:47:02] judged as worse than a hedge one but uh I I think yeah I don't think we're scoring things exactly right um and uh I'd say um it's actually like uh it's not clear exactly how to um okay but if you wanted to actually train with a let's say you wanted to train with something like a proper scoring function like you want that to be your reward uh like let's say we ask the model to Output probabilities on everything like uh it says 10 on this [00:47:34] sentence 20 on this sentence uh that would also have some problems because uh natural language is just very imprecise and that's what makes it powerful but uh it's um there's like uh just as much fuzziness on the sentence as whatever probability you're like depending on how you interpret the sentence like how you would make it there's some like underlying interpretation of it and but there's so many possible interpretations some would have low probabilities some would have high probability it's uh that [00:48:05] makes it very hard to do this um so yeah uh so I think this is a open problem maybe we should have some kind of formal statements of probability alongside the natural language statements but uh I don't know exactly how to do that or maybe we should uh like have some kind of uh set up some kind of objective where you have multiple agents collaborating and like they should Express uncertainty correctly because it's useful to the other agent or it's useful to itself later something like that [00:48:36] um so I'd say um okay so another class of open problems in this sort of general truthfulness direction is how do we go beyond um things that the labelers can easily do um so yeah it's just very hard to um uh to check a full a long answer about a technical subject or some Niche subject and uh so there's this general research area that's uh it's called scalable oversight in in the alignment it's sort [00:49:07] of in the alignment Community it's called scalable oversight but you could uh I'd say um like I say the ideas that uh um it's it's often easier to verify that something is that a solution is correct than to generate a correct solution right this is like a uh like a very basic like one of the most basic ideas in uh theoretical computer science so um and uh so you can uh like if you look at the P versus NP problem you could say [00:49:37] that uh one interpretation is that uh you can have a weak agent your verifier that provides an incentive to the strong agent so that uh the optimal when you optimize uh the strong agent you get you're solving a hard class of problems uh uh say like sat like sat is like the canonical problem that's like easy to check uh the solution but it's hard to find the assignment so you can uh but it so you can have a weak agent like that [00:50:07] only does a little bit of compute that provides the reward and uh that'll lead to uh solving a hard problem if you optimize your strong agent so yeah they're um so how do we so it seems like it should be possible to uh do some kind of uh it seems like it should be possible to have labelers to train a model to do things that are much too hard for the labelers to do themselves uh in in principle it should be possible to do this and uh [00:50:38] um so yeah there's a lot of ideas in this direction so uh there's um you can do things like um like you can try to delegate decompose the task a lot and delegate it like have your browsing model fact check each sentence and then like automatically aggregate all the results uh there's also the uh like you can also do some kind of mechanism design so that's more like getting it that's more like um this idea of setting up incentives so [00:51:08] you can set up some kind of game where you have competing agents that are competing for uh approval of your verifier uh and trying to one is saying why the other is wrong uh and uh there's a nice idea there about uh called AI safety via debate um Yeah so basically there's there's some work in this direction it's all pretty new and I think like we still have yet to see really good practical implementations of this stuff but it's starting to become necessary because uh [00:51:38] it's getting really hard for labelers to keep up with the models um yeah and last I would say like most speculatively I would say uh one unsatisfying thing about RL from Human feedback is this purely optimizing on human approval uh and uh um we don't always uh know the right answer and we're probably wrong about lots of things so it would be uh like uh like so we're just optimizing for what sounds convincing and what sounds right [00:52:08] uh what's kind of the knowledge of the day it would be easier it would be like it would be great if we could uh optimize for actual truth and uh have like somehow add more compute and have the models would train the models harder and have them get closer to the real uh truth so how do you do that uh so one idea is that if you have some kind of ground Truth uh you can uh you can optimize for actual uh like correctness so if you for example predicting the [00:52:38] future uh there's a million predictions about the future you can make and uh if you're um uh so if we use that as the reward function we might be able to uh like generate real knowledge and have a real test for the knowledge uh and uh yeah so that that's kind of prediction is one source of generating knowledge and you can also obviously do deduction uh you can if you have some kind of formal system or semi-formal reasoning system you can generate new knowledge by [00:53:09] deduction so I think getting our models to do that is another interesting challenge all right that's all thanks for your attention [applause] foreign [00:54:00] okay yeah yeah that would be a good one hey John you enter yeah so I'm wondering can you say can you say can you say anything about the the the the aspect of it it seems to have an element of creativity and that means where you give it say a patent or a pair of patents and you say put these together and come up with something new that in a new invention and it seems to do reasonably [00:54:30] well with that does that surprise you or how when you consider that new knowledge in a certain way oh yeah yeah I guess I could uh cons yeah that uh seems like it could be new knowledge I mean I guess uh there's some uh taste that you'd be injecting uh by asking it that question in the first place uh either that it's a good idea to combine inventions or that these are particular in uh like uh promising inventions to combine so uh so it's like you're collaborating with the model to create knowledge to some extent but uh yeah I would say uh yeah there's not a [00:55:01] fine line between creativity and uh just kind of uh learning pattern recognition pattern completion I think I'm on the same topic um uh so I've been I've been trading the models on classical literature philosophy um and I'm curious on a question say like what is beauty where there's no obvious fixed answer but there's many other answers I'm curious what how do you evaluate whether [00:55:31] if at all these quantitative kind of measurements of the relative Matrix of different answers about beauty I mean do they have any Pro you know precedence over an output oh yeah yeah I mean I talked about the difficulty of uh writing answers even when they're about uh like until even when there's no uh sort of um subjective even the if there's supposed to be objective and they're not like values [00:56:01] loaded or anything so yeah if you have something that's uh like uh is going to depend on taste and values then that's that's even much harder uh so uh yeah I don't think we have a good answer for that I mean the direction we've been going uh so far um just is to um like uh we don't think the model should have opinions on things yet so we want the model to instead be able to describe uh the set of opinions that humans have so I would want the [00:56:32] model to I would want the model to sort of um redirect that into a more factual question about what are some human theories about what are some what are the schools have thought that humans have on this John um uh yeah so first just like um to be really close yeah uh I just want to give you props for uh uh I think it might have been five or six years ago I participated in a in an AI progress forecasting meeting with John uh and uh [00:57:06] he was the only person who's for a math camp and he was the only person in the room more bullish than me on predicting AI progress um and I think he serves a lot of credit not only for building what he's built but for for having optimism years in advance that this kind of thing was possible so I just wanted to just uh call you out for that um and uh I wanted to ask about the um the web GPT demo um the it's really great how it gives this inner monologue and I'm wondering if if you have any what what's your [00:57:38] level of like optimism versus skeptimism skepticism for using that sort of you know inner monologue format for interpretability like can you distill a model so that it doesn't have enough room in its inner layers to think and it needs this inner monologue um and so we might be able to like read out its thoughts what are is this something you thought about it yeah oh yeah definitely uh so I would say um to the extent that we can't find uh Perfect Solutions for interpretability [00:58:08] or for making sure our models are safe or well-intentioned I think this is like a really good partial solution and we should do as much of it as possible um actually yeah yeah so I would say uh I would say it's very helpful for interpretability uh they're obviously like uh you can't completely trust it Like the Model could be uh trying to give it could be producing a deceptive inner monologue so that's definitely a concern but like you said you could also kind of use a small model so it has to [00:58:38] use the inner monologue to reach a certain level of intelligence of course then you could worry that it's doing some kind of steganography and it's hiding information but yeah it's a little far-fetched so overall I would say I think it's promising but maybe there's some theoretical concerns with it I also think uh like one thing I didn't mention is that if you have detailed inner monologues that allows you to use a shorter Horizon uh feedback so for example for browsing if you don't have the inner monologue and you see one [00:59:08] action like scroll you have no idea if this action makes sense or not so it's impossible to provide a reward on it but if the model says I'm scrolling to look for blah and then it has the scroll action a human can look at that single action and decide if it makes sense or not and uh so like by having inner monologue you can uh make the feedback at a you could train with RL at a shorter time Horizon and that also makes uh the system safer because you're not [00:59:39] optimizing for like long-term Behavior which could lead to weird results this art of choosing the smallest possible oh I'm sorry okay cool if you could say something about the uh if there's any recent work any current catch you later okay sorry sorry about that um hi John so I have a question about the uh so you mentioned in it earlier that there is some intrinsic Knowledge Graph in the models yeah and then you [01:00:09] showed an example of the of the model explaining dagger versus Fleet hacker right so dagger is able to directly explain it I assume that's because the knowledge is is in this inside of the model then it is still able to go into the web and search for Fleet uh dagger and be able to explain it but um I'll assume that that knowledge that that has some new Concepts that's not inside of the knowledge graph of the model so what do you expect uh the the difference in the capabilities of the models and then explain the two concepts [01:00:40] if there's any I didn't cast the lessons so like I guess if if the knowledge about dagger is inside and the knowledge about Fleet Decker is partly like outside of the models then do you expect any um difference in capabilities of models in explaining it to the concepts um yeah I guess uh um I would say there's probably some uh I'd say the model is probably best with Concepts that it has uh deeply uh like [01:01:10] that are deeply ingrained and it's seen them in a million contexts uh and if it's just seeing the concept for the first time in some document that it's conditioning on uh it's probably gonna have less intelligent things to say about it it's kind of like if you just uh this is just me kind of uh like half answer answering based on introspection or based on like I'm just kind of speculating here but I would expect that something that's like deeply ingrained uh would be easier to the model would be more intelligent at talking about that [01:01:41] so I would I would say would be better at talking about dagger than Fleet dagger for Fleet dagger it's going to have it's just going to say some kind of summary of uh what's in the document and it's not going to say anything too insightful about it okay thank you hey John we're going to make this the last question okay one more questions for the rest of your schedule tonight we want last one so you mentioned in part one the problem of the model learning to withhold information when that's not desirable uh [01:02:11] do you foresee that there could be issues with uh a conflict between the incentive of training the model not to withhold information and open domain contexts while also training it to not produce unsupported information in closed domain contexts even when it actually knows that information um yeah I think there's an extremely strong uh conflict between uh well there's a Precision recall kind of [01:02:41] uh conflict and there's uh there's um yeah there's a conflict between informativeness and uh correctness and and we I think you you often run into this when you're training so rlh like we're we're choosing some particular uh like reasonable point that we think is reasonable on this trade-off curve of how often the model should guess but it's unavoidable that there's a trade-off there all right let's thank John again [applause] foreign [01:03:15] thanks