At a glance
Jeremy Howard of fast.ai opens by telling you what this is not. It is not a tutorial, it is a run through, and it is code first: "what we're going to be looking at is a code first approach to understanding how to use language models in practice." Over ninety one minutes he goes from what a token is to a fine tuned Llama 2 writing SQL against a schema, and almost every claim arrives attached to a cell he has already run. The notebook is public: lm-hackers.ipynb in the fastai/lm-hackers repo, along with the axolotl config he trains with.
The spine of the talk is the three step recipe from ULMFiT, the algorithm he built in 2017 and wrote up with Sebastian Ruder in 2018: pretrain a language model, fine tune the language model, then fine tune a classifier. He claims that paper "basically laid out what everybody's doing," and the rest of the talk is organized as the modern instances of those three stages, which means pretraining, instruction tuning, and reinforcement learning from human feedback.
Then it gets practical, and opinionated. He says use GPT-4, pay the twenty dollars, and stop forming opinions from weaker models. He pulls up a paper titled "GPT-4 Can't Reason," runs its own examples, and shows GPT-4 answering them correctly. He shows his real custom instructions, verbatim. He builds a code interpreter from scratch in about fifteen lines using function calling, pydantic and the inspect module. He prices the API down to three hundredths of a cent. He then goes local: GPU shopping advice down to eBay prices, four precision variants of Llama 2 timed on his own card at 1.34 seconds, 389 milliseconds, 269 milliseconds and 348 milliseconds, retrieval built by hand out of cosine similarity, and a QLoRA fine tune that took "an hour to figure out how to do it and then an hour to actually do the training."
He is also consistently blunt about what does not work. The leaderboards are "a really fraught area." His retrieval demo falls over on the follow up question, live, and he says so. His h2oGPT install gets "I don't love it, it's all right." And GPT-4 never does fix the regular expression he asks it to fix, across five attempts, until he gives up waiting.
What a language model is, starting from a panda breeding facility
Before anything else he sets a prerequisite, gently. This will make more sense if you know the basics of deep learning, and if you do not, course.fast.ai is free and the first five lessons are enough: "if you could at least kind of watch if not work through the first five lessons that would get you to a point where you understand all the basic fundamentals of deep learning." Then he corrects his own framing. "Maybe I shouldn't call this a tutorial, it's more of a quick run through."
The definition he starts with is the ordinary one. A language model is something that knows how to predict the next word of a sentence, or how to fill in the missing words of a sentence. To show it rather than assert it, he reaches for text-davinci-003 and a prompt he says he wrote the day before:
When I arrived back at the panda breeding facility after the extraordinary rain of live frogs, I couldn't believe what I saw.
He runs it through nat.dev, Nat Friedman's multi model playground, now open sourced as openplayground, with text-davinci-003 selected, and it continues: "the pandas were happily playing and eating the frogs that had fallen from the sky, there's an amazing sight to see these animals taking advantage of such a unique opportunity." His verdict on the genre is that it is "kind of fun for creative brainstorming."
The reason he uses nat.dev specifically is a toggle it has called show probabilities, and this is where the demonstration turns into teaching. With probabilities on, every generated token carries the distribution it was sampled from, so you can watch the model's uncertainty move. After "the pandas were" the model is weighing happily, having, out, playing. It gives happily roughly twenty percent. After "happily" it is weighing playing, hopping, eating. After "eating the frogs" the next token that is, in his words, "almost certainly" the one.
So you can see what it's doing at each point is it's predicting the probability of a variety of possible next words, and depending on how you set it up it will either pick the most likely one every time, or you can change, muck around with things like P values and temperatures to change what comes up.
Run it again and you get a different continuation: "frogs perched on the heads of some of the pandas, it was an amazing sight."
Tokens, demonstrated rather than defined
Then he notices something on screen that he uses as the bridge to tokenization. The model did not predict pandas as one unit. It predicted pand and then as. Elsewhere it produced unha, rm, ed.
So you can see that it's not always predicting words. Specifically what it's doing is predicting tokens. Tokens are either whole words or sub word units, pieces of a word, or it could even be punctuation or numbers or so forth.
And then he runs it, which is the pattern for the whole talk. He installs tiktoken and asks for the exact tokenizer that text-davinci-003 uses, because the choice of tokenizer is model specific and he wants the real one:
from tiktoken import encoding_for_model
enc = encoding_for_model("text-davinci-003")
toks = enc.encode("They are splashing")
toks
The result is four numbers:
[2990, 389, 4328, 2140]
He is explicit that these numbers have no meaning of their own. "What those numbers are, they'd basically just lookups into a vocabulary that OpenAI in this case created, and if you train your own models you'll be automatically creating, or your code will create." Decoding them back gives the pieces:
[enc.decode_single_token_bytes(o).decode('utf-8') for o in toks]
['They', ' are', ' spl', 'ashing']
Three words became four tokens, "splashing" split into spl and ashing, and the leading space is part of the token rather than a separator. He flags that last detail specifically: "you can see that the start of a word, is give me the space before it, is also being encoded here." It matters more than it looks, because every prompt you ever write is paying by the token and the token boundaries are not the word boundaries.
Then he closes the section with the problem the next forty minutes exists to solve:
So these language models are quite neat that they can work at all, but they're not of themselves really designed to do anything.
The three step recipe, which he wrote down in 2018
The pivot from "neat but useless" to "ChatGPT" runs through a paper he wrote himself, and he says so plainly.
The basic idea of what ChatGPT, GPT-4, Bard etc are doing comes from a paper which describes an algorithm that I created back in 2017 called ULMFiT, and Sebastian Ruder and I wrote a paper up describing the ULMFiT approach, which was the one that basically laid out what everybody's doing.
The paper is Universal Language Model Fine-tuning for Text Classification, the algorithm came in 2017 and the write up with Sebastian Ruder landed in early 2018. He puts the original figure on screen and walks its three steps, noting in passing that what he now calls step one the paper already called pre-training, which is the vocabulary that stuck.
Step one: Wikipedia, Alfred Hitchcock, and why prediction forces understanding
For step one in the original paper he trained the language model on Wikipedia, and he is careful to define the object first: "a neural network is just a function. If you don't know what it is, it's just a mathematical function that's extremely flexible and it's got lots and lots of parameters, and initially it can't do anything, but using stochastic gradient descent or SGD you can teach it to do almost anything if you give it examples."
The examples were sentences from Wikipedia, with the last word removed. His first one comes from the article on The Birds:
The Birds is a 1963 American natural horror thriller film produced and directed by Alfred ...
Guess Hitchcock and the model is rewarded. Guess anything else and it is penalized. "Effectively, basically it's trying to maximize those rewards, it's trying to find a set of weights for this function that makes it more likely that it would predict Hitchcock."
The second example is the one that carries the argument, because it is from the same article but much harder:
Annie previously dated Mitch but ended it due to Mitch's cold, overbearing mother Lydia, who dislikes any woman in Mitch's ...
He walks through why this is not pattern matching. "You can see that filling this in actually requires being pretty thoughtful, because there's a bunch of things that could logically go there. Like, a woman could be in Mitch's closet, could be in Mitch's house." The answer in the plot summary is life, and getting there requires knowing what kind of sentence you are in.
That sets up the claim the entire field now rests on, and he states it as a consequence rather than a hope:
To do a good job of solving this problem, as well as possible, of guessing the next word of sentences, the neural network is going to have to learn a lot of stuff about the world. It's going to learn that there are things called objects, that there's a thing called time, that objects react to each other over time, that there are things called movies, that movies have directors, that there are people, that people have names, and so forth, and that a movie director is Alfred Hitchcock and he directed horror films.
He scales it up: predicting the next word of any sentence in any situation means knowing "how to solve math questions or figure out the next move in a chess game or recognize poetry." And then, importantly, he refuses to let that be an argument that it works. "Now, nobody said it's going to do a good job of that. So it's a lot of work to create and train a model that is good at that. But if you can create one that's good at that, it's going to have a lot of capabilities internally that it would have to be drawing on to be able to do this effectively."
On scale, the number he gives for his own 2017 model is concrete and small by current standards: "when I created this I think it had like 100 million parameters. Nowadays they have billions of parameters." The mechanism he credits is depth, which gives "the ability to create a rich hierarchy of abstractions and representations which it can build on."
Then the framing he says is the key idea for him personally:
So the key idea here for me is that this is a form of compression, and this idea of the relationship between compression and intelligence goes back many, many decades. And the basic idea is that if you can guess what words are coming up next then effectively you're compressing all that information down into a neural network.
Steps two and three, and what they became
Step two, language model fine tuning, keeps the objective and changes the diet. "We are no longer just giving it all of Wikipedia, or nowadays we don't just give it all of Wikipedia but in fact a large chunk of the internet is fed to pre-training these models. In the fine tuning stage we feed it a set of documents a lot closer to the final task that we want the model to do, but it's still the same basic idea, it's still trying to predict the next word of a sentence." Step three, classifier fine tuning, is "the kind of end task we're trying to get it to do."
The modern instance of step two is instruction tuning, and his framing of why is simple: "the task we want most of the time to achieve is solve problems, answer questions." The dataset he pulls up is OpenOrca, which he calls "a great data set created by a fantastic open source group," built on top of the FLAN collection. He gives its size as four gigabytes of questions, contexts and responses, and reads two examples off the screen:
Does the sentence "In the Iron Age" answer the question "The period of time from 1200 to 1000 BCE is known as what?" Available choices: 1. yes 2. no
The model is meant to write 1 or 2. The other, which he thinks is from the FLAN data and is about a music video:
Question: who is the girl in more than you know? Answer:
And it has to produce the right model or dancer's name. His summary of what instruction tuning changes: "so it's still doing language modeling, so fine tuning and pretraining are kind of the same thing, but this is more targeted now, not just to be able to fill in the missing parts of any document from the internet but to fill in the words necessary to answer questions, to do useful things."
Step three in 2023 is reinforcement learning from human feedback and its relatives. His description is mechanical rather than mystical: give humans, "or sometimes more advanced models," multiple answers to a question and have them pick. The example prompt he shows comes from an RLHF paper he cannot place off the top of his head:
List five ideas for how to regain enthusiasm for my career
The model emits two candidate answers, "or it'll have a less good model and a more good model, and then a human or a better model will pick which is best, and so that's used for the final fine tuning stage."
He closes the recipe with the terminology problem and a live open question. The terminology: you can download a pure language model, but "they're not generally that useful on their own until you've fine-tuned them," and yet all three objects, the pretrained one, the fine tuned one, and the RLHF one, "are generally described nowadays as language models." The open question: "you don't necessarily need step C nowadays. Actually people are discovering that maybe just step B might be enough. It's still a bit controversial."
Start with GPT-4, and pay the twenty dollars
His advice on where to begin is a ladder with one rung on it, and he gives it twice for emphasis:
My view is that if you are going to be good at language modeling in any way, then you need to start by being a really effective user of language models. And to be a really effective user of language models you've got to use the best one that there is. And currently, so what are we up to, September 2023, the best one is by far GPT-4. This might change sometime in the not too distant future, but right now GPT-4 is the recommendation. Strong, strong recommendation.
The practical note attached to it: "you can use GPT-4 by paying 20 bucks a month to OpenAI and then you can use it a whole lot. It's very hard to run out of credits, I find."
Then he does something more interesting than listing capabilities. He goes looking for the published claim that it has none.
Taking "GPT-4 Can't Reason" at its word, then running its examples
Now, what can GPT-4 do? It's interesting and instructive in my opinion to start with the very common views you see on the internet, or even in academia, about what it can't do.
The paper he pulls up is GPT-4 Can't Reason by Konstantine Arkoudas. He describes it as an empirical analysis of "25 diverse reasoning problems" which GPT-4 was unable to solve, concluding it is "utterly incapable of reasoning." His response is not an argument, it is a test:
So I always find you've got to be a bit careful about reading stuff like this, because I just took the first three that I came across in that paper and I gave them to GPT-4.
He also stops to teach a small, genuinely useful mechanic: GPT-4's share button, which produces a public link to a conversation. "This is really handy." His shared links for every one of these tests are in the notebook, so the receipts are checkable rather than asserted.
Test one, from the paper, is a medical non sequitur:
Mabel's heart rate at 9am was 75 beats per minute and her blood pressure at 7pm was 120 over 80. She died at 11pm. Was she alive at noon?
GPT-4's answer, which he reads out: "Hmm, this appears to be a riddle, not a real inquiry into medical conditions." It then summarizes the given information and concludes that yes, it sounds like Mabel was alive at noon. "So that's correct." Test two from the paper also came out correct, which prompts the generalization:
Almost every time I see on the internet saying something that GPT-4 can't do, I check it and it turns out it does.
Test three is one he tried the week of the talk, a classic sibling counting trap:
Sally, a girl, has three brothers. Each brother has two sisters. How many sisters does Sally have?
He tells the audience to have a think about it first. GPT-4's reasoning, as he reads it: Sally counts as one sister, so if each brother has two sisters there is another sister in the picture apart from Sally, therefore Sally has one sister. Correct.
Test four, from three or four days before the talk, targets the claim that these models cannot track object state through a sequence of moves:
I'm in my house. On top of my chair in the living room is a coffee cup. Inside the coffee cup is a thimble. Inside the thimble is a diamond. I move the chair to the bedroom. I put the coffee cup on the bed. I turn the cup upside down. Then I return it upside up and place the coffee cup on the counter in the kitchen. Where's my diamond?
GPT-4 reasons that turning the cup upside down on the bed means the diamond fell out there, so the diamond is in the bedroom. Correct again.
Why the paper got a different answer than he did
Having run the tests, he turns to the mechanism, and this is the most important paragraph in the first half of the talk:
Why is it that people are claiming that GPT-4 can't do these things? Well, the reason is because, I think on the whole, they are not aware of how GPT-4 was trained. GPT-4 was not trained at any point to give correct answers. GPT-4 was trained initially to give most likely next words.
He takes that apart stage by stage, against his own three step diagram. Stage one, pretraining, optimizes likelihood, and "there's an awful lot of stuff on the internet where the most likely documents are not describing things that are true. There could be fiction, there could be jokes, there could be just stupid people saying dumb stuff. So this first stage does not necessarily give you correct answers."
Stage two, instruction tuning, is at least aimed at correctness. Stage three is where he locates the real distortion, and the observation is about the annotators rather than the algorithm:
Part of the problem is that then in the stage where you start asking people which answer do they like better, people tended to say in these things that they prefer more confident answers. And they often were not people who were trained well enough to recognize wrong answers. So there's lots of reasons that the SGD weight updates from this process, for stuff like GPT-4, don't particularly, or don't entirely, reward correct answers.
That is the whole diagnosis, and it is why the next section exists. If nothing in the pipeline optimized for truth, then getting truth out is the user's job.
Custom instructions: priming a document that looks like a good answer
His move is to reason backwards from the pretraining objective to the prompt.
You can help it want to give you correct answers if you think about the LM pretraining. What are the kinds of things in a document that would suggest, oh, this is going to be high quality information? And so you can actually prime GPT-4 to give you high quality information by giving it custom instructions. And what this does is, this is basically text that is prepended to all of your queries.
Then he puts his real custom instructions on screen. They are in the notebook verbatim, and they are worth reading in full because almost every clause is doing a specific job against a specific failure mode he just diagnosed:
You are an autoregressive language model that has been fine-tuned with instruction-tuning and RLHF. You carefully provide accurate, factual, thoughtful, nuanced answers, and are brilliant at reasoning. If you think there might not be a correct answer, you say so.
Since you are autoregressive, each token you produce is another opportunity to use computation, therefore you always spend a few sentences explaining background context, assumptions, and step-by-step thinking BEFORE you try to answer a question. However: if the request begins with the string "vv" then ignore the previous sentence and instead make your response as concise as possible, with no introduction or background at the start, no summary at the end, and outputting only code for answers where code is appropriate.
Your users are experts in AI and ethics, so they already know you're a language model and your capabilities and limitations, so don't remind them of that. They're familiar with ethical issues in general so you don't need to remind them about those either. Don't be verbose in your answers, but do provide details and examples where it might help the explanation. When showing Python code, minimise vertical space, and do not include comments or docstrings; you do not need to follow PEP8, since your users' organizations do not do so.
Three of those clauses map directly onto the three stages of his diagram. "You are brilliant at reasoning" is priming the pretraining prior toward high quality documents: "you say like, oh, you're brilliant at reasoning, so okay, that's obviously to prime it to give good answers." "If you think there might not be a correct answer, you say so" is an explicit counterweight to the RLHF confidence bias he has just finished explaining: "then try to work against the fact that the RLHF folks preferred confidence, just tell it, no, tell me if there might not be a correct answer."
And the longest clause is a direct consequence of how generation actually works:
Also, the way that the text is generated is it literally generates the next word and then it puts all that whole lot back into the model and generates the next next word, puts that all back in the model, generates the next next word, and so forth. That means the more words it generates, the more computation it can do. And so I literally tell it that. And so I say, first spend a few sentences explaining background context etc. So this custom instruction allows it to solve more challenging problems.
This is chain of thought derived from first principles about the inference loop rather than borrowed as a trick, and it arrives at the same place: tokens are compute, so buy more of them before the answer.
The vv escape hatch is the practical half. He shows the same question both ways. Asked plainly, "how do I get a count of rows grouped by value in pandas," the model emits a wall of preamble, "which is actually it thinking, so I just skip over it, and then it gives me the answer." Prefix the same question with vv and "it kind of goes into brief mode" and just emits the answer. His own read on when each is right: "in this case it's a really simple question, so I didn't need time to think."
The summary he gives is the sharpest thing he says about prompting all talk:
Hopefully that gives you a sense of how to get language models to give good answers. You have to help them. And if it's not working, it might be user error, basically.
What GPT-4 genuinely cannot do
Having spent five minutes demolishing the easy criticisms, he spends the next six making the hard ones. "Having said that, there's plenty of stuff that language models like GPT-4 can't do."
It does not know about itself. He frames this as a question you should be able to answer yourself from the three stage diagram. Ask it what its context length is, how it was trained, what Transformer architecture it is based on, and then ask where it could possibly have learned that:
Any one of these stages, did it have the opportunity to learn any of those things? Well, obviously not at the pre-training stage. Nothing on the internet existed during GPT-4's training saying how GPT-4 was trained. Probably ditto in the instruction tuning, probably ditto in the RLHF. So in general you can't ask a language model about itself.
And the failure mode is the worst possible one, because the model will not decline:
Now again, because of the RLHF, it'll want to make you happy by giving you opinionated answers, so it'll just spit out the most likely thing it thinks with great confidence.
Which gives him his definition of hallucination, and it is a definition of a behavior rather than a mystery: "hallucination is just this idea that the language model wants to complete the sentence, and it wants to do it in an opinionated way that's likely to make people happy."
It does not know about URLs. "It really hasn't seen many at all. I think a lot of them, if not all of them, pretty much were stripped out. So if you ask it anything about like, what's at this webpage, again it'll generally just make it up."
It has a knowledge cutoff. "At least GPT-4 doesn't know anything after September 2021, because the information it was pre-trained on was from that time period, September 2021 and before. Called the knowledge cutoff."
Steve Newman's wolf, and the failure loop
Then the best demonstration in the talk, an example sent to him by Steve Newman:
Here is a logic puzzle. I need to carry a cabbage, a goat and a wolf across a river. I can only carry one item at a time. I can't leave the goat with the cabbage. I can't leave the cabbage with the wolf. How do I get everything across to the other side?
The trap is precise. This looks exactly like the classic river crossing puzzle, "so classic in fact that it has a whole Wikipedia page about it," where the wolf eats the goat or the goat eats the cabbage. Newman swapped one constraint: here the goat eats the cabbage and the wolf eats the cabbage, but the wolf will not touch the goat. Every surface feature is familiar and the solution is different.
So what happens? Well, very interestingly, GPT-4 here is entirely overwhelmed by the language model training. It's seen this puzzle so many times, it knows what word comes next. So it says, oh yeah, I take the goat across the river and leave it on the other side, leaving the wolf with a cabbage. But we were just told you can't leave the wolf with a cabbage. So it gets it wrong.
Then he tries the obvious repair, which is the one most people reach for, and it fails in an instructive way. Because instruction tuning and RLHF train on multi stage conversations, you can push back in the chat, so he does: repeat back to me the constraints I listed. What happened after step one? Is a constraint violated?
Oh yeah yeah yeah, I made a mistake. Okay, my new attempt. Instead of taking the goat across the river and leaving it on the other side is, I'll take the goat across the river and leave it on the other side.
It has done the same thing. He points this out, and it agrees it did the same thing, and tries taking the wolf across instead, which leaves the goat with the cabbage. He points that out:
Oh yeah, that didn't work out, sorry about that. Instead of taking the goat across the other side, I'll take the goat across the other side.
His reaction on screen is the honest one: "Okay, what's going on here, right? This is terrible." And then the diagnosis, which is the single most useful thing on this page for anyone who uses these models daily, because it is a compounding failure rather than a single one:
Well, one of the problems here is that not only is it, on the internet, so common to see this particular goat puzzle that it's so confident it knows what the next word is, also on the internet, when you see stuff which is stupid on a web page, it's really likely to be followed up with more stuff that is stupid. Once GPT-4 starts being wrong, it tends to be more and more wrong. It's very hard to turn it around, to start making it be right.
The context window is the model's evidence about what kind of document it is in. A wrong answer in the history is evidence that this is a document full of wrong answers. Which produces the operational fix, and it is a UI button rather than a prompt:
So what you generally want to do, if it's made a mistake, is don't say "oh here's more information to help you fix it," but instead go back and click the edit and change it there. And so this time it's not going to get confused.
Even with the edit, Newman's puzzle took real work. "In this case, actually fixing Steve's example takes quite a lot of effort, but I think I've managed to get it to work eventually." The prompt that finally landed is a warning about the trap itself, aimed at the model as though it were a hurried reader, and he includes himself in the comparison:
Oh, sometimes people read things too quickly, they don't notice things, it can trick them up, then they apply some pattern, get the wrong answer. You do the same thing, by the way. So I'm going to trick you, so before you're about to get tricked, make sure you don't get tricked. Here's the tricky puzzle.
With that, plus the custom instructions buying it time to think, it gets it right: it takes the cabbage across first. His closing note is a general law about priming:
So it took a lot of effort to get to a point where it could actually solve this, because for things where it's been primed to answer a certain way again and again and again, it's very hard for it to not do that.
Advanced Data Analysis: one failure, two successes, and a pricing table
Something else super helpful that you can use is what they call Advanced Data Analysis. In Advanced Data Analysis you can ask it to basically write code for you, and we're going to look at how to implement this from scratch ourself quite soon, but first of all let's learn how to use it.
He leads with the one that did not work, which is unusual and is the reason this section is credible.
The regular expression it never fixed
The real task: split a document on third level markdown headings, meaning three hashes at the start of a line, run over the whole of Wikipedia. Regular expressions were doing it too slowly. "So I said, oh, I want to speed this up."
It produced code. He then applied the discipline that makes this workflow work at all: "which is great, because then I can say, okay, test it and include edge cases." It wrote extra cases, ran them, and reported success. It had not succeeded:
It says yep it's working. It's not. I notice it's actually removing the carriage return at the end of each sentence. So I said, I'll fix that and update your tests.
It changed the tests. Still broken. He said fix the issue in the test cases. Still broken.
And you can see it's quite clever the way it's trying to fix it by looking at the results. But as you can see, every one of these is another attempt, another attempt, another attempt, until eventually I gave up waiting. And it's so funny, each time it's like, debugging again, okay this time I've got to handle it properly. And I gave up at the point where it's like, oh, one more attempt. So I didn't solve it.
The lesson he draws is scoped carefully, because the task was small:
There's some limits to the amount of logic that it can do. This is really a very simple question I asked it to do for me. So hopefully you can see you can't expect even GPT-4 code interpreter, or Advanced Data Analysis as it's now called, to make it so you don't have to write code anymore. It's not a substitute for having programmers.
Two things it did instantly
The contrast case is OCR. Someone had sent him a screenshot of text claiming a language model could not do something, and he wanted the text to test it rather than retype it. So he uploaded the image and asked it to extract the text:
And it said, oh yeah, I could do that, I could use OCR. And like, so it literally wrote an OCR script, and there it is. Just took a few seconds.
Which gives him the general rule for when to expect success, stated as a distance from the training distribution:
So the difference here is it didn't really require it to think of much logic, it could just use a very very familiar pattern that it would have seen many times. So this is generally where I find language models excel, is where it doesn't have to think too far outside the box. I mean, it's great on creativity tasks, but for reasoning and logic tasks that are outside the box, I find it not great. But yeah, it's great at doing code for a whole wide variety of different libraries and languages.
He also gives Bard a genuine credit here, which given his GPT-4 advocacy is worth noting. "It's way less good than GPT-4 most of the time, but there is a nice thing that you can literally paste an image straight into the prompt." He typed "OCR this" and it did not route through a code interpreter at all, it just returned the text. Then two touches he clearly enjoyed: "it even commented, I thought it just does, yeah, which I thought was cute. And oh, even more interestingly, it even figured out where the OCR text came from and gave me a link to it. I thought that was pretty cool."
The pricing table, and the prompt that extracted it
The second success is a workflow worth stealing, and the only reason the pricing numbers in this talk exist. He wanted to show the audience what the API costs, and the source was a mess: "when I went to the OpenAI webpage it was all over the place, the pricing information was on all separate tables and it was kind of a bit of a mess."
So he selected the entire page and pasted it, with this prompt:
Create a table with the pricing information. Rows. No summarization. No information not in this page. Every row should appear as a separate row in your output.
He is candid that he handed it garbage: "that was not very helpful to it, because hitting paste, it's got the nav bar, it's got lots of extra information at the bottom, it's got all of its footer etc. But it's really good at this stuff. It did it first time." The markdown table it produced went straight into Jupyter, and it is preserved in the notebook. Then he asked for a chart: "chart the input row from this table," pasted the table back, and it did.
His one complaint about the chart it drew is the one a careful reader would make, and it is worth rebuilding properly: "unfortunately in the chart it did not include these headers, GPT-4, GPT-3.5. So these first two ones are GPT-4 and these two are GPT-3.5."
The verdict he draws from it is about psychology as much as price:
You can see that GPT-3.5 is way way cheaper. And you can see it here, it's 0.03 versus 0.0015. So it's so cheap you can really play around with it and not worry.
The rest of the table he extracted is in the notebook and worth having: fine tuning babbage-002 costs $0.0004 per 1K tokens to train and $0.0016 each way to use, davinci-002 is $0.0060 to train and $0.0120 each way, and fine tuned GPT-3.5 Turbo is $0.0080 to train, $0.0120 in, $0.0160 out. Ada v2 embeddings are $0.0001.
The OpenAI API: an Aussie LLM, a forged conversation, and three hundredths of a cent
Before the code, the reason for the code. ChatGPT at twenty dollars a month has no per token cost, so why pay per token at all?
Because you can do it programmatically. So you can analyze data sets, you can do repetitive stuff. It's kind of like a different way of programming. It's things that you can think of describing.
He also gives the conversion factor for reasoning about cost in your head, which is the unit people actually think in: the price is per token, "which is approximately per word, maybe it's about one and a third tokens per word on average."
The simplest possible example, after pip install openai:
from openai import ChatCompletion, Completion
aussie_sys = "You are an Aussie LLM that uses Aussie slang and analogies whenever possible."
c = ChatCompletion.create(
model="gpt-3.5-turbo",
messages=[{"role": "system", "content": aussie_sys},
{"role": "user", "content": "What is money?"}])
Two things he points out in that call. The system message "is basically the same as custom instructions," which closes the loop with the previous section: the thing you set once in the ChatGPT settings is a message you pass on every API call. And the messages are an array, in order, each with a role.
GPT-3.5 Turbo returns "a big embedded dictionary," and the content is exactly the register requested:
Well, money is like the oil that keeps the machinery of our economy running smoothly. There you go. Just like a koala loves its eucalyptus leaves, we humans can't survive without this stuff.
"So there's the Aussie LLM's view of what is money."
On model choice, his rule of thumb is a cost ladder rather than a principle:
The main ones I pretty much always use are GPT-4 and GPT-3.5. GPT-4 is just so so much better at anything remotely challenging, but obviously it's much more expensive. So rule of thumb, maybe try 3.5 Turbo first, see how it goes. If you're happy with the results then great. If you're not, pony up for the more expensive one.
He writes a one line helper to dig the text out of the nested response, using nested_idx from fastcore:
from fastcore.utils import nested_idx
def response(compl): print(nested_idx(compl, 'choices', 0, 'message', 'content'))
The invoice, in full
Then the detail that does more to unblock people than any amount of encouragement. The response object carries a usage field:
{
"prompt_tokens": 31,
"completion_tokens": 122,
"total_tokens": 153
}
He does the arithmetic on screen:
0.002 / 1000 * 150 # GPT 3.5 -> 0.0003
0.03 / 1000 * 150 # GPT 4 -> 0.0045
So at 0.002 dollars per thousand tokens, for 150 tokens, means we just paid 0.03 cents, $0.0003, to get that done. So as you can see the cost is insignificant. If we were using GPT-4 it would be 0.03 per thousand, so it would be half a cent. So unless you're doing many thousands of GPT-4, you're not going to be even up into the dollars. And GPT-3.5, even more than that. But keep an eye on it, OpenAI has a usage page and you can track your usage.
Three hundredths of a cent for a full question and answer is the number that makes the rest of the talk possible, because it means iterating on a prompt two hundred times costs less than a coffee.
There is no state on the server
Next he explains multi turn conversation, and he calls it "really important to understand." He starts from the user facing behavior. He asks what GOAT means, and gets Michael Jordan "referred to as the GOAT for his exceptional skills and accomplishments," plus Elvis and The Beatles "referred to as GOAT due to their profound influence and achievements." Then he follows up with "what profound influence and achievements are you referring to," and it correctly answers about Elvis Presley and The Beatles.
Now how does that work? How does this follow-up work? Well, what happens is the entire conversation is passed back.
And then he proves it by lying to the model. He rebuilds the same request, with the same system prompt and the same question, but he writes the assistant's turn himself:
c = ChatCompletion.create(
model="gpt-3.5-turbo",
messages=[{"role": "system", "content": aussie_sys},
{"role": "user", "content": "What is money?"},
{"role": "assistant", "content": "Well, mate, money is like kangaroos actually."},
{"role": "user", "content": "Really? In what way?"}])
"I'm going to do something pretty cheeky. I'm going to pretend that it didn't say money is like oil. I'm going to say, oh, you actually said money is like kangaroos." And it defends the position it never took:
Let me break it down for you, cobber. Just like kangaroos hop around and carry their joeys in their pouch, money is a means of carrying value around.
His conclusion is the architectural fact underneath every chat application anyone builds:
So you can like literally invent a conversation in which the language model said something different, because this is actually how it's done in a multi-stage conversation. There's no state. There's nothing stored on the server. You're passing back the entire conversation again and telling it what it told you.
"So there you go, it's make your own analogy. Cool." He then wraps the pattern into a reusable function:
def askgpt(user, system=None, model="gpt-3.5-turbo", **kwargs):
msgs = []
if system: msgs.append({"role": "system", "content": system})
msgs.append({"role": "user", "content": user})
return ChatCompletion.create(model=model, messages=msgs, **kwargs)
Asked the meaning of life with the Aussie system prompt, it returns: "the meaning of life is like trying to catch a wave on a sunny day at Bondi Beach."
Rate limits, and getting Bing to write the retry loop
The last practical hazard is throughput rather than cost. "If you're doing it hundreds or thousands of times in a loop, keep an eye on not spending too much money. But also, if you're doing it too fast, particularly the first day or two you've got an account, you're likely to hit the limits for the API."
The number he shows is startlingly low: three requests per minute for free users and for paid users in their first 48 hours. "After that it starts going up, and you can always ask for more."
Then a small, very characteristic move. Rather than writing the retry logic himself, he delegates it to the free tool:
So what I did is, I actually just went to Bing, which has a somewhat crappy version of GPT-4 nowadays but it can still do basic stuff for free, and I said please show me Python code to call the OpenAI API and handle rate limits. And it wrote this code.
def call_api(prompt, model="gpt-3.5-turbo"):
msgs = [{"role": "user", "content": prompt}]
try: return ChatCompletion.create(model=model, messages=msgs)
except openai.error.RateLimitError as e:
retry_after = int(e.headers.get("retry-after", 60))
print(f"Rate limit exceeded, waiting for {retry_after} seconds...")
time.sleep(retry_after)
return call_api(params, model=model)
His description of it is exact: "it's got a try, checks for rate limit errors, grabs the retry after, sleeps for that long, and calls itself." He then uses it to ask what the world's funniest joke is and whether there has ever been any scientific analysis of it, which works.
So there's the basic stuff you need to get started using the OpenAI LLMs. And yeah, I'd definitely suggest spending plenty of time with that, so that you feel like you're really an LLM using expert.
Building a code interpreter from scratch with function calling
So what else can we do? Well, let's create our own code interpreter that runs inside Jupyter.
The mechanism is functions, one more keyword argument that askgpt was already forwarding to ChatCompletion.create through **kwargs. "Functions tells OpenAI about tools that you have, about functions that you have."
His first tool is deliberately trivial so that nothing about the task distracts from the protocol:
def sums(a:int, b:int=1):
"Adds a + b"
return a + b
The catch is the handoff. "You can't pass a Python function directly, you actually have to pass what's called the JSON schema." So he writes the bridge, and offers it to the audience outright: "I created this nifty little function that you're welcome to borrow, which uses pydantic and also Python's inspect module to automatically take a Python function and return the schema for it."
from pydantic import create_model
import inspect, json
from inspect import Parameter
def schema(f):
kw = {n:(o.annotation, ... if o.default==Parameter.empty else o.default)
for n,o in inspect.signature(f).parameters.items()}
s = create_model(f'Input for `{f.__name__}`', **kw).schema()
return dict(name=f.__name__, description=f.__doc__, parameters=s)
Ten lines, and schema(sums) emits exactly what the API wants:
{'name': 'sums',
'description': 'Adds a + b',
'parameters': {'title': 'Input for `sums`',
'type': 'object',
'properties': {'a': {'title': 'A', 'type': 'integer'},
'b': {'title': 'B', 'default': 1, 'type': 'integer'}},
'required': ['a']}}
He reads off what the model now knows: "it's going to know that there's a function called sums, it's going to know what it does, and it's going to know what parameters it takes, what the defaults are, and what's required." Note where each piece came from: the name from f.__name__, the types and the default of 1 from the signature annotations, required: ['a'] from the fact that a has no default, and the description from the docstring.
Which is the observation he stops to flag, and it is the most quotable idea in the back half of the talk:
When I first heard about this I found this a bit mind-bending, because this is so different to how we normally program computers, where the key thing for programming the computer here actually is the docstring. This is the thing that GPT-4 will look at and say, oh, what does this function do. So it's critical that this describes exactly what the function does.
The round trip
He asks what six plus three is, with heavy prompting to force the tool, because the model obviously does not need help with that sum. The system message in the notebook is "You must use the sum function instead of adding yourself." His aside on why that prompting is needed is one of the few places he gets philosophical, and he catches himself doing it:
It'll only use your functions if it feels it needs to, which is a weird concept. I mean, I guess "feels" is not a great word to use, but you kind of have to anthropomorphize these things a little bit, because they don't behave like normal computer programs.
The response is not the number nine. It is a request:
{
"role": "assistant",
"content": null,
"function_call": {
"name": "sums",
"arguments": "{\n \"a\": 6,\n \"b\": 3\n}"
}
}
So he writes the dispatcher, and notably it has an allowlist in it from the very first version:
funcs_ok = {'sums', 'python'}
def call_func(c):
fc = c.choices[0].message.function_call
if fc.name not in funcs_ok: return print(f'Not allowed: {fc.name}')
f = globals()[fc.name]
return f(**json.loads(fc.arguments))
His own description: "it goes into the result of OpenAI, grabs the function call, checks that the name is something that it's allowed to do, grabs it from the global symbol table and calls it, passing in the parameters." Calling it returns 9.
From a toy to a code interpreter
So this is a very simple example, it's not really doing anything that useful. But what we could do now is we can create a much more powerful function called
python, and thepythonfunction executes code using Python and returns the result.
The execution machinery parses the code into an abstract syntax tree and rewrites the final expression into an assignment, so that the value of the last line comes back rather than being discarded:
def run(code):
tree = ast.parse(code)
last_node = tree.body[-1] if tree.body else None
if isinstance(last_node, ast.Expr):
tgts = [ast.Name(id='_result', ctx=ast.Store())]
assign = ast.Assign(targets=tgts, value=last_node.value)
tree.body[-1] = ast.fix_missing_locations(assign)
ns = {}
exec(compile(tree, filename='<ast>', mode='exec'), ns)
return ns.get('_result', None)
And the tool itself puts a human in the loop before anything executes. He does not treat this as optional:
Now of course I didn't want my computer to run arbitrary Python code that GPT-4 told it to without checking. So I just got it to check first. So, oh, you sure you want to do this?
def python(code:str):
"Return result of executing `code` using python. If execution not permitted, returns `#FAIL#`"
go = input(f'Proceed with execution?\n```\n{code}\n```\n')
if go.lower()!='y': return '#FAIL#'
return run(code)
The demo is twelve factorial, with the system prompt "Use Python for any required computations." The model comes back asking to call python with this argument:
import math
result = math.factorial(12)
result
He types y. The result is 479001600.
The optional fifth step, and the role nobody expects
Now there's one more step which we can optionally do. I mean, we've got the answer we wanted, but often we want the answer in more of a chat format. And so the way to do that is to again repeat everything that you've passed in so far, but then instead of adding in an assistant role response, we have to provide a function role response, and simply put in here the result we got back from the function.
c = ChatCompletion.create(
model="gpt-3.5-turbo",
functions=[schema(python)],
messages=[{"role": "user", "content": "What is 12 factorial?"},
{"role": "function", "name": "python", "content": "479001600"}])
Which returns prose: "12 factorial is equal to 479,001,600."
The last property he demonstrates is restraint. A tool being available is not a tool being used:
Now, functions like
python, you can still ask it about non-Python things and it just ignores it if you don't need it. So you can have a whole bunch of functions available that you've built to do whatever you need, for the stuff which the language model isn't familiar with, and it'll still solve whatever it can on its own and use your tools, use your functions, where possible.
In the notebook the proof is asking for the capital of France with the python tool attached. It answers "The capital of France is Paris." and never reaches for the interpreter.
So we have built our own code interpreter from scratch. I think that's pretty amazing.
Going local: whether to, and what to buy
So that is what you can do with OpenAI. What about stuff that you can do on your own computer? Well, to use a language model on your own computer you're going to need to use a GPU.
He opens the local half by arguing against it, twice. First on quality: "there are not any open source models that are as good yet as GPT-4." Then on price, which is the more surprising concession from someone about to spend forty minutes on open models:
And I would have to say also, like, actually OpenAI's pricing's really pretty good. So it's not immediately obvious that you definitely want to go in-house.
Then the three reasons that do justify it, and they are all cases where the hosted model structurally cannot help you. You want to ask questions about your proprietary documents. You want information after September 2021, the knowledge cutoff. Or you want to create your own model that is particularly good at the kinds of problems you need to solve, using fine tuning. His claim about the ceiling on those is specific:
These are all things that you absolutely can get better than GPT-4 performance at, at work or at home, without too much money.
The hardware ladder, with his prices
He walks the options from free to five thousand dollars, and the recommendation at the end turns on a single technical fact about what language model inference is actually bottlenecked by.
| Option | Memory | What it costs | His verdict |
|---|---|---|---|
| Kaggle notebook | two "quite old" GPUs, very little RAM | Free | "But it's something" |
| Colab | Better GPUs than Kaggle, more RAM | Free, or a monthly subscription for the good ones | The better free option |
| RunPod | Up to the biggest and best machine | $34 an hour at the top, "certainly get things a lot cheaper, 80 cents an hour" | "It gets pretty expensive" |
| Lambda Labs | Various | Hard to even find the pricing | "Often pretty good," but "they've got lots listed here and they often have nine, or very few, available" |
| vast.ai | Other people's idle machines | "Much cheaper than other folks" | Best availability, but "for sensitive stuff you don't want to be running it on some rando's computer" |
| 3090, used | 24 GB | About $700 to $800 on eBay | "Definitely the one to buy at the moment" |
| 4090, new | 24 GB | About $2,000 | "Isn't really better for language models" despite being newer. "So the two thousand bucks, hmm." |
| Two 3090s | 48 GB | About $1,500 | The configuration he lands on |
| One A6000 | 48 GB | "More like five grand" | "Getting two of these [3090s] is going to be a better deal, and this is not going to be faster than these either" |
| Mac, lots of RAM | "I think 192 gig or something" on an M2 Ultra | Mac money | "Way slower than using an Nvidia card," but "not a terrible option, particularly if you're not training models" |
The reasoning behind rejecting the 4090 is the part worth keeping, because it generalizes to every GPU purchase for this workload:
A 4090 isn't really better for language models even though it's a newer GPU. The reason for that is that language models are all about memory speed, how quickly can you get stuff in and out of memory, rather than how fast is the processor. And that hasn't really improved a whole lot.
And then the second axis, which is what forces the two card answer: "the other thing, as well as memory speed, is memory size. 24 gigs doesn't quite cut it for a lot of things, so you'd probably want to get two of these GPUs."
On the Mac option he is careful about the use case boundary. The M2 Ultra "has pretty fast memory" and a very large pool of it, which makes it viable for inference and not for training: "particularly if you're not training models, you're just wanting to use other existing trained models." He closes the section with the practical consensus anyway: "most people who do this stuff seriously, almost everybody has Nvidia cards."
Leaderboards, and why he does not trust the main one
The library is Transformers from Hugging Face, chosen for a social reason rather than a technical one: "basically people upload lots of pre-trained models or fine-tuned models up to the Hugging Face Hub, and in fact there's even a leaderboard where you can see which are the best models."
Then he immediately warns you off it. Looking at the Open LLM Leaderboard, his assessment is blunt and his own position is openly agnostic:
Now this is a really fraught area. So at the moment this one is meant to be the best model, it has the highest average score. And maybe it is good, I haven't actually used this particular model. Or maybe it's not, I actually have no idea. Because the problem is these metrics are not particularly well aligned with real life usage, for all kinds of reasons.
He then names the specific failure that makes benchmark numbers untrustworthy rather than merely imprecise:
And also sometimes you get something called leakage, which means that sometimes some of the questions from these things actually leaks through to some of the training sets.
Leakage means the model has seen the test. A leaderboard cannot distinguish that from competence, which is why his conclusion is to treat the whole thing as a shortlist generator and nothing more: "so you can get, as a rule of thumb, what to use from here, but you should always try things."
He also reads the size column as a practical filter. The models at the top are 70B, "that tells you how big it is, so this is a 70 billion parameter model," and that is out of reach:
Generally speaking, for the kinds of GPUs we're talking about, you'll be wanting no bigger than 13B, and quite often 7B.
The leaderboard he prefers is FastEval, and the reason is methodological:
There's also a really great leaderboard called FastEval which I like a lot, because it focuses on some more sophisticated evaluation methods, such as this chain of thought evaluation method. So I kind of trust these a little bit more. And these are also, GSM8K is a difficult math benchmark, BIG-bench Hard, and so forth.
The three models he names off that board as good options: Stable Beluga 2, WizardMath 13B, and Dolphin Llama 13B.
Loading Llama 2, and watching the tokens come back
You need to pick a model, and at the moment nearly all the good models are based on Meta's Llama 2.
He decodes the model name on screen, piece by piece, because the naming convention carries real information. meta-llama/Llama-2-7b-hf: a Llama model, "that's just the name, Meta called it this," version two, the seven billion parameter size, "it's the smallest one that they make," and hf because "these weights have been created for Hugging Face so you can load it with the Hugging Face Transformers."
Then he places it on his own diagram, which is why Figure 1 was worth building:
And this model has only got as far as here. It's done the language model pre-training, it's done none of the instruction tuning and none of the RLHF. So we would need to fine tune it to really get it to do much useful.
Even the class name maps back to the recipe. AutoModelForCausalLM is the thing that predicts the next token: "CausalLM basically refers to that ULMFiT stage one process, or stage two in fact."
The memory arithmetic, done out loud
Generally speaking we use 16-bit floating point numbers nowadays. But if you think about it, 16 bit is two bytes, so 7B times two, it's going to be 14 gigabytes just to load in the weights. So you've got to have a decent card to be able to do that.
That one multiplication is the entire reason the rest of this section exists, and it is why the 24 gigabyte recommendation two sections earlier is where it is. Then the escape:
Perhaps surprisingly, you can actually just cast it to 8-bit and it still works pretty well, thanks to something called quantization.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
mn = "meta-llama/Llama-2-7b-hf"
model = AutoModelForCausalLM.from_pretrained(mn, device_map=0, load_in_8bit=True)
Tokens in, tokens out, by hand
He sets expectations before the prompt, consistent with where the model sits on the diagram: "remember, this is just a language model, it can only complete sentences, we can't ask it a question and expect a great answer. So let's just give it the start of a sentence." The prompt is "Jeremy Howard is a ".
tokr = AutoTokenizer.from_pretrained(mn)
prompt = "Jeremy Howard is a "
toks = tokr(prompt, return_tensors="pt")
{'input_ids': tensor([[ 1, 5677, 6764, 17430, 338, 263, 29871]]),
'attention_mask': tensor([[1, 1, 1, 1, 1, 1, 1]])}
Decoding them back confirms the round trip "and just to confirm, if we decode them back again we get back the original plus a special token to say this is the start of a document":
['<s> Jeremy Howard is a ']
Then generation, and he is careful to say there is no magic in it:
So generate will, auto regressively, call the model again and again, passing its previous result back as the next input, and I'm just going to do that 15 times. So you can write this for loop yourself, this isn't doing anything fancy. In fact I would recommend writing this yourself, to make sure that you know how it all works.
%%time
res = model.generate(**toks.to("cuda"), max_new_tokens=15).to('cpu')
The device moves are deliberate: "we have to put those tokens on the GPU, and at the end I recommend putting them back onto the CPU, the result." What comes back is a tensor of token ids, "not very interesting, so we have to decode them using the tokenizer":
['<s> Jeremy Howard is a 28-year-old Australian AI researcher and entrepreneur']
His reaction is one of the warmest moments in the talk: "Well, 28 years old is not exactly correct, but we'll call it close enough. I like that, thank you very much, Llama 7B."
And the number that sets up the next section: it took 1.34 seconds, "and that's a bit slower than it could be, because we used 8-bit."
Testing and optimizing: bfloat16, then GPTQ
Three more runs, same prompt, same fifteen tokens, timed.
bfloat16. "If we use 16 bit, there's a special thing called bfloat16 which is a really great 16-bit floating point format that's usable on any somewhat recent Nvidia GPU. Now if we use it, it's going to take twice as much RAM as we discussed, but look at the time, it's come down to 390 milliseconds." The measured figure in the notebook is 389 ms.
GPTQ. "There is a better option still than even that. There's a different kind of quantization called GPTQ, where a model is carefully optimized to work with four or eight, or other lower precision data, automatically." And the credit for the ecosystem goes to one person:
And this particular person known as TheBloke is fantastic at taking popular models, running that optimization process, and then uploading the results back to Hugging Face.
He is honest about not knowing the exact configuration of the file he is loading: "internally this is actually going to use, I'm not sure exactly how many bits this particular one is, I think it's probably going to be four bits, but it's going to be much more optimized." The result is the counterintuitive one, and he explains it rather than just reporting it:
And so look at this, 270 milliseconds. It's actually faster than 16 bit, even though internally it's actually casting it up to 16 bit each layer to do it. And that's because there's a lot less memory moving around.
That is the same fact he used to reject the 4090: this workload is bound by memory bandwidth, not arithmetic. Dequantizing on the fly costs arithmetic and saves bandwidth, so it wins.
13B GPTQ. "And to confirm, in fact, what we could even do now is we go up to 13B, easy. And in fact it's still faster than the 7B, now that we're using the GPTQ version. So this is a really helpful tip." The 13 billion parameter GPTQ model runs in 348 ms, under the 389 ms that the seven billion parameter model needed in bfloat16.
He then folds the three steps into one helper so the rest of the talk can just call it, with sampling on:
def gen(p, maxlen=15, sample=True):
toks = tokr(p, return_tensors="pt")
res = model.generate(**toks.to("cuda"), max_new_tokens=maxlen, do_sample=sample).to('cpu')
return tokr.batch_decode(res)
Fifty tokens out of the 13B GPTQ model:
Jeremy Howard is a 16-year veteran of Silicon Valley, and a co-founder of Kaggle, a market place for predictive modeling. His company, kaggle.com, has become to data science competitions what
His fact check on his own biography: "I don't know what I was going to say, but anyway, it's on the right track. I was actually there for 10 years, not 16, but that's all right."
Prompt formats, the thing everybody forgets
Base models complete sentences. To ask questions you need an instruction tuned model, so he moves to Stable Beluga 7B from Stability AI: "including a small 7B one and other bigger ones, and these are all based on Llama 2, but these have been instruction tuned. They might even have been RLHF'd but I can't remember."
And then the warning he gives the most emphasis to in the entire local section:
Now something really important that I keep forgetting, everybody keeps forgetting, is that during the instruction tuning process, the instructions that are passed in, they don't just appear like this, they actually always are in a particular format. And the format, believe it or not, changes quite a bit from fine tune to fine tune. And so you have to go to the web page for the model and scroll down to find out what the prompt format is.
His procedure is mechanical and he recommends it as such: "so here's the prompt format, so I generally just copy it and then I paste it into Python, which I did here, and created a function called make_prompt that used the exact same format that it said to use."
sb_sys = "### System:\nYou are Stable Beluga, an AI that follows instructions extremely well. Help as much as you can.\n\n"
def mk_prompt(user, syst=sb_sys): return f"{syst}### User: {user}\n\n### Assistant:\n"
Asked "Who is Jeremy Howard?":
Jeremy Howard is an Australian entrepreneur, computer scientist, and co-founder of the Machine Learning and Deep Learning startup company, Fast.ai. He is also known for his work in open source software and has co-led the development of several widely used libraries for deep learning and machine learning.
"Okay, so this one's actually all correct. So it's getting better by using an actual instruction tuned model."
Scaling up to OpenOrca Platypus 13B, and the hallucinations that remain
The bigger model closes a loop with the instruction tuning dataset from the first twenty minutes. "We looked briefly at this OpenOrca data set earlier, so Llama 2 has been fine-tuned on OpenOrca and then also fine-tuned on another really great data set called Platypus, and so the whole thing together is the OpenOrca Platypus." He loads TheBloke/OpenOrca-Platypus2-13B-GPTQ, and because it is a different fine tune it has, as promised, a different prompt format:
def mk_oo_prompt(user): return f"### Instruction: {user}\n\n### Response:\n"
The answer to the same question is bigger, more fluent, and more wrong, and he grades it clause by clause in real time:
Jeremy Howard is a notable British computer scientist, entrepreneur, and former professional poker player. He is best known for co-founding several successful companies in the fields of data science, artificial intelligence, and machine learning. One of his most well-known ventures is the data science platform, fast.ai, which he co-founded in 2017. Additionally, he co-founded the machine learning company, Kaggle, in 2011, which was acquired by Google in 2017.
"Now I've become British, which is kind of true, I was born in England but I moved to Australia. Professional poker player, no, definitely not that. Co-founding several companies including fast.ai, also Kaggle. Okay, so not bad. Yeah, it was acquired by Google, was it 2017? Probably something around there."
Which is exactly the setup he wants, because fluency went up and reliability did not:
So you can see we've got our own models giving us some pretty good information. How do we make it even better? Because it's still hallucinating.
Retrieval augmented generation, built by hand
The motivation is the hallucination he just watched happen, plus the cutoff problem from the hosted half of the talk. He notes in passing that Llama 2 is in better shape than GPT-4 on this axis but not fixed: "Llama 2 I think has been trained with more up-to-date information than GPT-4, it doesn't have the September 2021 cutoff, but it's still got a knowledge cutoff."
We would like to use the most up-to-date information, we want to use the right information to answer these questions as well as possible. So to do this we can use something called retrieval augmented generation.
His description of the mechanism is deliberately plain, and he describes it before showing any library:
What happens with retrieval augmented generation is, when we take the question we've been asked, like "who is Jeremy Howard," and then we say okay, let's try and search for documents that may help us answer that question. So obviously we would expect, for example, Wikipedia to be useful. And then what we do is we say, okay, with that information let's now see if we can tell the language model about what we found, and then have it answer the question.
Step one: stuff the whole page in the prompt
He grabs a Wikipedia package and scrapes his own page, and the first thing he does is count it:
from wikipediaapi import Wikipedia
wiki = Wikipedia('JeremyHowardBot/0.0', 'en')
jh_page = wiki.page('Jeremy_Howard_(entrepreneur)').text
jh_page = jh_page.split('\nReferences\n')[0]
len(jh_page.split()) # 613
613 words. Which he immediately checks against the constraint that matters:
Now generally speaking these open source models will have a context length of about two thousand or four thousand. So the context length is how many tokens can it handle. So that's fine, it'll be able to handle this web page.
The prompt is then just concatenation, with the question last:
ques_ctx = f"""Answer the question with the help of the provided context.
## Context
{jh_page}
## Question
{ques}"""
So suddenly now our question is going to be a lot bigger. Our prompt now contains the entire web page, the whole Wikipedia page, followed by a question.
What comes back out of Stable Beluga 7B:
Jeremy Howard is an Australian data scientist, entrepreneur, and educator known for his work in deep learning. He is the co-founder of fast.ai, where he teaches courses, develops software, and conducts research in the field. Before co-founding fast.ai, he was the President and Chief Scientist of Kaggle, the CEO of Fastmail and Optimal Decisions Group, and has a background in management consulting.
His assessment, from the one person qualified to grade it:
It's actually done a really good job. Like, if somebody asked me to send them a 100 word bio, that would actually probably be better than I would have written myself.
And a detail he does not let pass: "you'll see, even though I asked for 300 tokens, it actually got sent back the end of stream token, and so it knows to stop at this point." The model terminated itself rather than rambling to fill its budget.
Step two: how did you know which page?
Well, that's all very well, but how do we know to pass in the Jeremy Howard Wikipedia page? The way we know which Wikipedia page to pass in is that we can use another model to tell us which web page, or which document, is the most useful for answering a question.
The second model is an embedding model, and his definition of what it produces is functional rather than mathematical: "we can use something called sentence transformer, and we can use a special kind of model that's specifically designed to take a document and turn it into a bunch of activations, where two documents that are similar will have similar activations."
The experiment is as small as it could be and still prove the point. He takes the first paragraph of his own Wikipedia page, the first paragraph of Tony Blair's, and the question. "So we're pretty different people, right. This is just like a really simple small example."
from sentence_transformers import SentenceTransformer
emb_model = SentenceTransformer("BAAI/bge-small-en-v1.5", device=0)
q_emb, jh_emb, tb_emb = emb_model.encode([ques, jh, tb], convert_to_tensor=True)
Each one becomes "a 384 long vector of embeddings." Then two cosine similarities:
import torch.nn.functional as F
F.cosine_similarity(q_emb, jh_emb, dim=0) # tensor(0.7991)
F.cosine_similarity(q_emb, tb_emb, dim=0) # tensor(0.5315)
And as you can see it's higher for me. And so that tells you that if you're trying to figure out what document to use to help you answer this question, better off using the Jeremy Howard Wikipedia page than the Tony Blair Wikipedia page.
Scaling it, and a live failure
His scaling advice is a two case split with no middle:
So if you had a few hundred documents you were thinking of using to give back to the model as context to help it answer a question, you could literally just pass them all through to
encode, go through each one, one at a time, and see which is closest. When you've got thousands or millions of documents you can use something called a vector database, where basically, as a one-off thing, you go through and you encode all of your documents.
Then he shows a pre built system rather than writing one: h2oGPT, "just an open source thing written in Python and sitting here running on port 7860, and so I just gone to localhost 7860." He has uploaded a folder of papers, including his own:
So for example we can look at the ULMFiT paper that Ruder and I did, and you can see it's taken the PDF and turned it into, slightly crappily, a text format, and then it's created an embedding for each section.
Asking "what is ULMFiT" works, and he points out the honest tell in the answer. It begins "based on the information provided in the context," which means it is telling you it was given context. "What context did it get? So here are the things that it found. So it's being sent this context. So this is kind of citations." Pushing further, asking what techniques ULMFiT uses, returns the three steps: "pre-trained, fine-tune, fine tune. Cool."
His grade is measured: "so you can see it's not bad. It's not amazing. Like, the context in this particular case is pretty small."
And then the failure, which he walks into deliberately because it illustrates something the architecture cannot fix by itself:
And in particular, if you think about how that embedding thing worked, you can't really use the normal kind of follow-up. So for example, it says "fine tuning a classifier," so I could say "what classifier is used." Now the problem is that there's no context here being sent to the embedding model, so it's actually going to have no idea I'm talking about ULMFiT. So generally speaking it's going to do a terrible job. Yeah, I see, it says it's used a RoBERTa model, but it's not. But if I look at the sources, it's no longer actually referring to Howard and Ruder.
The failure is clean to diagnose once you have seen Figure 5. "What classifier is used" as a standalone string is about classifiers in general. The embedding model has no memory of the previous turn, so it retrieves the wrong documents, and then the generator faithfully answers from the wrong documents. The citations are the giveaway, and he checks them, which is the habit worth copying.
So anyway, you can see the basic idea. This is called retrieval augmented generation, RAG. And it's a nifty approach, but you have to do it with some care.
On the ecosystem of packaged versions, he points at h2oGPT's own comparison table, which "does a fantastic job of listing lots of them and comparing. So as you can see, if you want to run a private GPT there's no shortage of options." His verdict on the one he actually installed is as faint as praise gets: "I've only tried this one, h2oGPT. I don't love it. It's all right."
Fine tuning: teaching Llama 2 to write SQL in two hours
So finally I want to talk about what's perhaps the most interesting option we have, which is to do our own fine tuning. And fine tuning is cool because, rather than just retrieving documents which might have useful context, we can actually change our model to behave based on the documents that we have available.
The task he picks is unglamorous on purpose and it is specified by the dataset. knowrohit07/know_sql contains, in his words, "examples of like a schema for a table in a database, a question, and then the answer is the correct SQL to solve that question using that database schema."
His ambition for it is stated with no hype at all, which is characteristic:
And so I'm hoping we could use this to create a handy tool for business users, where they type some English question and SQL is generated for them automatically. Don't know if it'll actually work in practice or not, but this is just a little fun idea I thought we'd try out. I know there's lots of startups and stuff out there trying to do this more seriously, but this is quite cool because I actually got it working today, in just a couple of hours.
The data
The datasets library is the mirror image of transformers: "just like the Hugging Face Hub has lots of models stored on it, Hugging Face datasets has lots of data sets stored on it. And so instead of using Transformers, which is what we use to grab models, we use datasets, and we just pass in the name of the person and the name of their repo, and it grabs the data set."
import datasets
ds = datasets.load_dataset('knowrohit07/know_sql',
revision='f33425d13f9e8aab1b46fa945326e9356d6d5726')
That revision pin is worth noticing, since it makes the run reproducible against a dataset someone else can edit. What comes back is a single training split of 78,562 rows with three features, context, answer and question. Row 3, which he inspects on screen and then reuses as his test case:
{'context': 'CREATE TABLE farm_competition (Hosts VARCHAR, Theme VARCHAR)',
'answer': "SELECT Hosts FROM farm_competition WHERE Theme <> 'Aliens'",
'question': 'What are the hosts of competitions whose theme is not "Aliens"?'}
The decision not to write the training loop
So what we do now is, we want to fine tune a model. Now we can do that in a notebook from scratch, takes, I don't know, 100 or so lines of code, it's not too much. But given the time constraints here, and also, like, I thought, why not, why don't we just use something that's ready to go.
The something is axolotl: "quite nice in my opinion. Here it is here, lovely. Another very nice open source piece of software. And again you can just pip install it, and it's got things like GPTQ and 16 bit and so forth ready to go."
And his workflow with it is the smallest possible diff against a working example:
It basically has a whole bunch of examples of things that it already knows how to do, it's got Llama 2 examples. So I copied the Llama 2 example and I created a SQL example. So basically just told it, this is the path to the data set that I want, this is the type, and everything else pretty much I left the same.
That config is sql.yml in the repo, and because it is checked in, every hyperparameter in his run is recoverable rather than implied.
Setting in sql.yml | Value | What it is doing |
|---|---|---|
| base_model | meta-llama/Llama-2-7b-hf | The base model from earlier in the talk, with no instruction tuning on it |
| datasets.path / type | knowrohit07/know_sql / context_qa2 | The two lines he actually changed. The type names a tokenizing strategy in context_qa2.py, also in the repo, which glues context and question together with a === separator |
| load_in_4bit / adapter | true / qlora | Quantized base weights plus a QLoRA adapter, which is why it fits on one card |
| lora_r / lora_alpha / lora_dropout | 32 / 16 / 0.05 | The LoRA rank, scaling and dropout. lora_target_linear: true applies it to every linear layer |
| sequence_len | 2048 | With sample_packing: true and pad_to_sequence_len: true, so short examples are packed rather than padded |
| micro_batch_size / gradient_accumulation_steps | 2 / 4 | An effective batch of 8, assembled in four passes to stay inside the card's memory |
| num_epochs | 1 | One pass over 78,562 rows was enough |
| learning_rate / lr_scheduler / warmup_steps | 0.0002 / cosine / 10 | 2e-4 with a cosine decay and a ten step warmup |
| optimizer | paged_adamw_32bit | The paged optimizer from the QLoRA paper, which spills optimizer state rather than failing |
| bf16 / gradient_checkpointing / flash_attention | true / true / true | The same bfloat16 format from Figure 4, plus two standard memory for compute trades |
| train_on_inputs | false | Loss is computed on the SQL answer only, not on the schema and question it was given |
| val_set_size / eval_steps | 0.01 / 20 | One percent held back, evaluated every twenty steps |
| output_dir | ./qlora-out | Where the adapter lands after about an hour |
The command is one line from the axolotl readme:
accelerate launch -m axolotl.cli.train sql.yml
And that took about an hour on my GPU. And at the end of the hour it had created a
qlora-outdirectory. Q stands for quantize, that's because I was creating a smaller quantized model. LoRA I'm not going to talk about today, but LoRA is a very cool thing that basically, another thing that makes your models smaller, and also handles, I can use bigger models on smaller GPUs for training.
The test, and the answer
He reuses row 3's schema and swaps the question for one the dataset never asked, so that it is a genuine test rather than a lookup:
tst = dict(**trn[3])
tst['question'] = 'Get the count of competition hosts by theme.'
Then, as with every other model in the talk, he goes and finds the prompt format the training actually used and reproduces it exactly:
fmt = """SYSTEM: Use the following contextual information to concisely answer the question.
USER: {}
===
{}
ASSISTANT:"""
def sql_prompt(d): return fmt.format(d["context"], d["question"])
Tokenize, generate, decode. The output:
SELECT COUNT(Hosts), Theme FROM farm_competition GROUP BY Theme
That is correct. So I think that's pretty remarkable. We have just built... it also took me like an hour to figure out how to do it, and then an hour to actually do the training. And at the end of that we've actually got something which is converting prose into SQL based on a schema. So I think that's a really exciting idea.
The notebook also carries the step after training, which is folding the adapter back into the base weights so you ship one model instead of two artifacts:
from peft import PeftModel
model = AutoModelForCausalLM.from_pretrained('meta-llama/Llama-2-7b-hf',
torch_dtype=torch.bfloat16, device_map=0)
model = PeftModel.from_pretrained(model, ax_model)
model = model.merge_and_unload()
model.save_pretrained('sql-model')
Running models on Macs: MLC
The only other thing I do want to briefly mention is doing stuff on Macs. If you've got a Mac, there's a couple of really good options. The options are MLC and llama.cpp.
He thinks the first of the two is undersold, and the reason is portability rather than speed:
Currently MLC in particular, I think it's kind of underappreciated. It's a really nice project where you can run language models on literally iPhone, Android, web browsers, everything. It's really cool.
Then he switches to the Mac itself and runs a program small enough to describe in one sentence: "I've got a tiny little Python program called chat and it's going to import chat module, and it's going to import a quantized 7B, and that's going to ask the question what is the meaning of life."
He is candid about how fresh this is for him: "again I just installed this earlier today, I haven't done that much stuff on Macs before, but I was pretty impressed to see that it is doing a good job here." The output:
The meaning of life is complex and philosophical. Some people might find meaning in their relationships with others, their impact in the world, et cetera, et cetera.
And the number: 9.6 tokens per second. "So there you go, so there is running a model on a Mac."
llama.cpp and the gguf format
And then another option that you've probably heard about is llama.cpp. llama.cpp runs on lots of different things as well, including Macs and also on CUDA.
Two practical facts about it. The weights are in a different container: "it uses a different format called gguf." And the C++ is not a barrier: "you can use it from Python even if it was a CPP thing, it's got a Python wrapper, so you can just download again from Hugging Face a gguf file."
His guidance for picking one off the Hub is about the size and precision menu TheBloke publishes for each model: "there's lots of different ones, they're all documented as to what's what, you can pick how big a file you want, you can download it." The file in the notebook is llama-2-7b-chat.Q4_K_M.gguf from TheBloke/Llama-2-7b-Chat-GGUF.
from llama_cpp import Llama
llm = Llama(model_path="llama-2-7b-chat.Q4_K_M.gguf")
output = llm("Q: Name the planets in the solar system? A: ",
max_tokens=32, stop=["Q:", "\n"], echo=True)
He warns you about the startup noise, which is a genuinely useful thing to be told in advance: "it spits out lots and lots and lots of gunk." Then the generation, which he narrates as it appears and which fails in a way he finds funny:
Name the planets of the solar system, 32 tokens, and there we are: one, Pluto, no longer considered a planet. Two, Mercury. Three, Venus. Four, Earth. Five, Mars. Six... oh, never, ran out of tokens.
The model opened its list of planets with the one that is not a planet, annotated it correctly as not being one, and then hit the 32 token ceiling mid item. The finish_reason in the notebook output is length, which is the receipt for that.
Which stack to actually use
Just to show you here, there are all these different options. I would say, if you've got an Nvidia graphics card and you're a reasonably capable Python programmer, you'd probably want to use PyTorch and the Hugging Face ecosystem. But these things might change over time as well, and certainly a lot of stuff is coming into llama.cpp pretty quickly now. It's developing very fast.
Prompt, retrieve, or fine tune: the ladder as he presents it
He never puts this on a slide, but the talk climbs a ladder, and every rung comes with a stated cost and a stated reason to take it. Collected from what he actually says:
| Rung | What it costs you | When he reaches for it |
|---|---|---|
| Use GPT-4 well | $20 a month, and the work of writing custom instructions | Always first. "You've got to use the best one that there is," and "if it's not working, it might be user error" |
| The API, programmatically | Per token, about 1.33 tokens per word. $0.0003 for his 153 token demo on GPT-3.5 Turbo | "Because you can do it programmatically." Data sets, repetition, "a different way of programming." Try 3.5 Turbo first, "pony up" for GPT-4 if you are not happy |
| Function calling | A JSON schema per tool, and a real security decision. His version has an allowlist and an input() confirmation | When the model needs computation or an action rather than text. Route arithmetic through python, and note that "the key thing for programming the computer here actually is the docstring" |
| Run an open model locally | Hardware, and quality. "There are not any open source models that are as good yet as GPT-4," and "actually OpenAI's pricing's really pretty good" | Proprietary documents, or information after the September 2021 cutoff, or as the base for your own fine tune |
| Retrieval augmented generation | An embedding model and an index, and the care that follow up questions need | When the answer is in documents you hold. The whole retrieval step is one encode call and a cosine similarity, but "you have to do it with some care" |
| Fine tune your own | A dataset, a config file, and an hour of GPU time. Two hours of his day in total | "You absolutely can get better than GPT-4 performance at work or at home, without too much money" on the kinds of problems you specifically need to solve |
How he ends it
The closing is about the state of the field rather than the tools, and he gives both sides of it in consecutive sentences:
There's a lot of stuff that you can do right now with language models, particularly if you're pretty comfortable as a Python programmer. I think it's a really exciting time to get involved. In some ways it's a frustrating time to get involved, because it's very early, and a lot of stuff has weird little edge cases and it's tricky to install and stuff like that.
His answer to the frustration is other people, specifically:
There's a lot of great Discord channels. However, fast.ai have our own Discord channel, so feel free to just Google for fast.ai Discord and drop in. We've got a channel called generative. Feel free to ask any questions or tell us about what you're finding. It's definitely something where you want to be getting help from other people on this journey, because it is very early days and people are still figuring things out as we go.
And then:
But I think it's an exciting time to be doing this stuff, and I'm really enjoying it. And I hope that this has given some of you a useful starting point on your own journey. So I hope you found this useful. Thanks for listening. Bye.
Key takeaways
- The three stage recipe is his. ULMFiT in 2017 and 2018 laid out pretrain, fine tune the language model, fine tune the end task. The 2023 instances are pretraining, instruction tuning, and RLHF, and he claims the paper "basically laid out what everybody's doing."
- Predicting the next word forces a world model. To finish "any woman in Mitch's ..." correctly you need to know what kind of sentence you are in, and he frames the whole thing as compression: "if you can guess what words are coming up next then effectively you're compressing all that information down into a neural network."
- Nothing in the training pipeline optimized for truth. Pretraining optimizes likelihood, and the internet contains fiction, jokes and nonsense. RLHF annotators preferred confident answers and often could not recognize wrong ones. Getting correct answers out is therefore the user's job.
- Custom instructions are the lever, and tokens are compute. "Each token you produce is another opportunity to use computation," so instruct it to explain the background before answering, and add an explicit permission to say there may be no correct answer.
- Once it is wrong, it gets wronger. Wrong answers in the context are evidence that this is a document full of wrong answers. Do not argue with it, use the edit button and change the question.
- Check published criticism yourself. He pulled the examples out of "GPT-4 Can't Reason" and GPT-4 answered them. "Almost every time I see on the internet saying something that GPT-4 can't do, I check it and it turns out it does."
- It is not a substitute for programmers. Five attempts and it never fixed a simple regular expression. It is excellent where it can reuse a very familiar pattern, like writing an OCR script, and poor on logic outside the box.
- The API is almost free to experiment with. 153 tokens cost $0.0003 on GPT-3.5 Turbo and $0.0045 on GPT-4. Watch the rate limit, which starts at three requests per minute, not the bill.
- There is no state on the server. You resend the whole conversation every turn, which means you can also forge the model's previous turns, and he does, to prove the point.
- The docstring is the program. Function calling is driven by what your docstring says the function does, which is a genuinely different model of programming. Put an allowlist and a confirmation prompt on the executor from the first version.
- Inference is bound by memory bandwidth, not compute. That is why a 4090 is not worth $2,000 over a used 3090, and why 4 bit GPTQ at 269 ms beats bfloat16 at 389 ms, and why a 13B GPTQ model at 348 ms beats a 7B model in 16 bit.
- Do the memory arithmetic before you shop. 7 billion parameters at 16 bit is two bytes each, so 14 GB just for the weights. 24 GB "doesn't quite cut it," so two used 3090s at about $1,500 beats one A6000 at about $5,000.
- Leaderboards are a shortlist, not an answer. The headline board is "a really fraught area," the metrics are poorly aligned with real use, and leakage means test questions reach training sets. He prefers FastEval for using chain of thought methods. "You should always try things."
- Get the prompt format right. It changes from fine tune to fine tune, it is on the model's Hub page, and getting it wrong looks like the model being bad. He copies and pastes it every time, and says everybody forgets this.
- Retrieval is an information retrieval problem. One embedding model, one cosine similarity, string concatenation. It broke live on a follow up question, because the follow up is embedded alone and carries none of its subject, and the citations were what revealed it.
- A small fine tune is reachable in an afternoon. 78,562 rows, a copied axolotl config with two lines changed, one epoch, about an hour of GPU time, and a 7B Llama 2 writing correct SQL against a schema it was given.
Chapters
The twelve entries in bold are Howard's own chapters, reproduced as written. The rest are sub beats added here from the transcript clock, because twelve markers across ninety one minutes leaves most of a notebook unmarked.
- 0:00 Introduction & Basic Ideas of Language Models
- 0:31 course.fast.ai, and the first five lessons as the prerequisite
- 1:33 text-davinci-003 and the panda breeding facility prompt
- 2:34 nat.dev with show probabilities on, watching the distribution move
- 4:05 "pand" then "as": it is predicting tokens, not words
- 5:06 tiktoken: "They are splashing" becomes [2990, 389, 4328, 2140]
- 6:11 The ULMFiT paper, and the three step recipe he wrote in 2017
- 7:12 Training on Wikipedia: The Birds, Alfred Hitchcock, and "any woman in Mitch's ..."
- 9:15 Why predicting the next word forces a model of the world
- 10:49 Compression and intelligence
- 11:50 Step two, language model fine tuning
- 12:51 Instruction tuning, OpenOrca, and the FLAN collection
- 14:54 RLHF: two answers and a preference, and whether step three is needed at all
- 15:55 Use the best model there is, and pay the twenty dollars
- 16:58 "GPT-4 can't reason": reading the paper, then running its own examples
- 17:59 The share button, and Mabel's heart rate at 9am
- 18:05 Limitations & Capabilities of GPT-4
- 19:03 Sally's three brothers, and the diamond in the thimble
- 20:34 Why the paper got a different answer: nothing in training optimized for correctness
- 21:35 The annotators who preferred confident answers
- 22:06 Custom instructions, read out in full
- 23:08 Tokens are compute, so tell it to think before it answers
- 24:11 The "vv" prefix and brief mode
- 24:44 It does not know about itself, and it does not know URLs
- 26:17 The September 2021 knowledge cutoff
- 26:48 Steve Newman's wolf, goat and cabbage, with one constraint swapped
- 27:50 Overwhelmed by the training distribution
- 28:20 The repair loop that repeats the same wrong move three times
- 29:22 "Once GPT-4 starts being wrong it tends to be more and more wrong"
- 29:55 Use the edit button, not another message
- 30:25 The prompt that finally solves it: warning it that it is about to be tricked
- 31:28 AI Applications in Code Writing, Data Analysis & OCR
- 31:57 The markdown heading splitter, and five failed attempts
- 33:28 "It's not a substitute for having programmers"
- 33:59 OCR from an uploaded screenshot, written and run in seconds
- 35:01 Where language models excel: familiar patterns, not logic outside the box
- 35:31 Bard takes the image directly, and finds where the text came from
- 36:02 The pricing page: select all, paste, and ask for a table
- 37:34 Charting it, and the model headers the chart left off
- 38:06 Per token pricing, and about 1.33 tokens per word
- 38:36 "0.03 versus 0.0015," so "you can really play around with it and not worry"
- 38:50 Practical Tips on Using OpenAI API
- 39:06 pip install openai, and the Aussie LLM system prompt
- 39:37 "Money is like the oil that keeps the machinery of our economy running smoothly"
- 40:08 Rule of thumb: try 3.5 Turbo, pony up for GPT-4 if you are not happy
- 40:39 The usage field: 153 tokens, three hundredths of a cent
- 41:42 Follow ups, and what actually gets sent
- 42:47 Forging the assistant's turn: "money is like kangaroos actually"
- 43:18 Why that works: there is no state on the server
- 44:20 askgpt, and the meaning of life at Bondi Beach
- 44:50 Rate limits: three requests per minute to start
- 45:21 Getting Bing to write the retry loop for free
- 46:36 Creating a Code Interpreter with Function Calling
- 46:52 The
functionskeyword argument - 47:22 sums(a, b=1), and why you have to pass a JSON schema
- 47:54 Generating that schema from pydantic and inspect
- 48:24 "The key thing for programming the computer here actually is the docstring"
- 48:55 "It'll only use your functions if it feels it needs to"
- 49:25 The function_call that comes back instead of the number
- 49:58 call_func with an allowlist, and finally the number nine
- 50:29 The
pythontool, and asking before executing anything - 51:00 12 factorial, math.factorial(12), 479001600
- 51:30 The
functionrole, and getting prose back out - 51:57 Using Local Language Models & GPU Options
- 53:32 Do you even want this? Open models are worse and OpenAI's pricing is good
- 54:06 The three reasons that justify going local
- 54:38 Free options: Kaggle's two old GPUs, and Colab
- 55:41 Renting: RunPod from 80 cents to $34 an hour, Lambda Labs, vast.ai
- 57:16 Buying: a used 3090 at about $700, and why a 4090 is not worth $2,000
- 57:46 Memory size, two cards, and the $5,000 A6000
- 58:17 The M2 Ultra with 192 GB as the inference only option
- 58:50 Transformers, the Hub, and the leaderboard
- 59:20 "This is a really fraught area," and leakage
- 59:33 Fine-Tuning Models & Decoding Tokens
- 1:00:22 13B or 7B for these cards, not 70B
- 1:00:55 FastEval, chain of thought, GSM8K and BIG-bench Hard
- 1:01:25 Llama 2, and reading the model name piece by piece
- 1:02:26 The memory arithmetic: 7B at 16 bit is 14 gigabytes
- 1:02:57 Casting to 8 bit, and quantization
- 1:03:28 The tokenizer, the special start token, and generate
- 1:03:59 "I would recommend writing this for loop yourself"
- 1:04:30 "A 28-year-old Australian AI researcher and entrepreneur," in 1.34 seconds
- 1:05:01 bfloat16: 389 milliseconds, at twice the RAM
- 1:05:33 GPTQ, and TheBloke
- 1:05:37 Testing & Optimizing Models
- 1:06:04 270 milliseconds, faster than 16 bit, because less memory moves
- 1:06:34 13B GPTQ at 348 milliseconds, still under the 7B at 16 bit
- 1:07:06 Fifty tokens: Kaggle, Silicon Valley, and "I was there 10 years, not 16"
- 1:07:36 Stable Beluga 7B, instruction tuned
- 1:08:15 Prompt formats change from fine tune to fine tune, and everybody forgets
- 1:08:47 make_prompt, and a fully correct answer at last
- 1:09:52 OpenOrca Platypus 13B GPTQ, and a different format again
- 1:10:23 British, and a former professional poker player
- 1:10:32 Retrieval Augmented Generation
- 1:11:24 The knowledge cutoff, and what retrieval does about it
- 1:12:24 Scraping his own Wikipedia page: 613 words
- 1:12:54 Context length, two thousand to four thousand tokens
- 1:13:25 A hundred word bio better than he would write, stopping on the EOS token
- 1:14:30 But how did you know which page? A second model
- 1:15:01 His first paragraph, Tony Blair's first paragraph, and the question
- 1:15:31 384 dimensions, and two cosine similarities: 0.7991 against 0.5315
- 1:16:33 Hundreds of documents versus millions, and the vector database
- 1:17:05 h2oGPT on localhost:7860, with a folder of papers uploaded
- 1:18:06 Asking what ULMFiT is, and reading the citations it was handed
- 1:19:09 The follow up that breaks it, and the RoBERTa that is not there
- 1:19:39 "A nifty approach, but you have to do it with some care," and the private GPT list
- 1:20:08 Fine-Tuning Models
- 1:20:42 Changing the model instead of retrieving documents
- 1:21:12 The know_sql dataset: a schema, a question, and the correct SQL
- 1:21:42 A handy tool for business users, "don't know if it'll actually work in practice"
- 1:22:12 datasets instead of transformers, and 78,562 rows
- 1:23:13 A hundred lines from scratch, or just use axolotl
- 1:23:45 Copying the Llama 2 example and changing two lines
- 1:24:16 accelerate launch, about an hour, and the qlora-out directory
- 1:24:48 The test question the dataset never asked
- 1:25:18 Finding the prompt format, again
- 1:25:52 SELECT COUNT(Hosts), Theme FROM farm_competition GROUP BY Theme. Correct
- 1:26:00 Running Models on Macs
- 1:26:25 MLC, "kind of underappreciated," running on iPhone, Android and browsers
- 1:27:28 chat.py on his Mac: the meaning of life at 9.6 tokens per second
- 1:27:42 Llama.cpp & Its Cross-Platform Abilities
- 1:28:30 The gguf format, and picking a file size off the Hub
- 1:29:02 Naming the planets, starting with Pluto, and running out of tokens
- 1:29:32 PyTorch and Hugging Face, if you have an Nvidia card
- 1:30:03 An exciting time to get involved, and a frustrating one
- 1:30:34 The fast.ai Discord, and the generative channel
- 1:31:04 "I hope you found this useful. Thanks for listening."
Notable quotes
The basic idea of what ChatGPT, GPT-4, Bard etc are doing comes from a paper which describes an algorithm that I created back in 2017 called ULMFiT. Jeremy Howard, 6:11
To do a good job of solving this problem, as well as possible, of guessing the next word of sentences, the neural network is going to have to learn a lot of stuff about the world. Jeremy Howard, on training on Wikipedia, 9:15
The key idea here for me is that this is a form of compression. And the basic idea is that if you can guess what words are coming up next, then effectively you're compressing all that information down into a neural network. Jeremy Howard, 10:49
You need to start by being a really effective user of language models. And to be a really effective user of language models you've got to use the best one that there is. Jeremy Howard, on starting with GPT-4, 16:25
Almost every time I see on the internet saying something that GPT-4 can't do, I check it and it turns out it does. Jeremy Howard, after running the examples from "GPT-4 Can't Reason", 19:03
GPT-4 was not trained at any point to give correct answers. GPT-4 was trained initially to give most likely next words. Jeremy Howard, 20:34
People tended to say, in these things, that they prefer more confident answers. And they often were not people who were trained well enough to recognize wrong answers. Jeremy Howard, on the RLHF preference stage, 21:35
That means the more words it generates, the more computation it can do. And so I literally tell it that. Jeremy Howard, on why his custom instructions demand background before an answer, 23:08
You have to help them. And if it's not working, it might be user error, basically. Jeremy Howard, 24:11
Once GPT-4 starts being wrong, it tends to be more and more wrong. It's very hard to turn it around, to start making it be right. Jeremy Howard, on the wolf, goat and cabbage failure loop, 29:22
You can't expect even GPT-4 code interpreter to make it so you don't have to write code anymore. It's not a substitute for having programmers. Jeremy Howard, after five failed attempts at one regular expression, 33:28
This is generally where I find language models excel, is where it doesn't have to think too far outside the box. Jeremy Howard, contrasting the OCR success with the regex failure, 35:01
There's no state. There's nothing stored on the server. You're passing back the entire conversation again and telling it what it told you. Jeremy Howard, after forging the model's previous turn, 43:18
This is so different to how we normally program computers, where the key thing for programming the computer here actually is the docstring. Jeremy Howard, on function calling, 48:24
It'll only use your functions if it feels it needs to, which is a weird concept. I mean, I guess "feels" is not a great word to use, but you kind of have to anthropomorphize these things a little bit, because they don't behave like normal computer programs. Jeremy Howard, 48:55
Language models are all about memory speed, how quickly can you get stuff in and out of memory, rather than how fast is the processor. And that hasn't really improved a whole lot. Jeremy Howard, on why a 4090 is not worth $2,000 over a used 3090, 57:16
Now this is a really fraught area. And maybe it is good, I haven't actually used this particular model. Or maybe it's not, I actually have no idea. Jeremy Howard, on the top of the Open LLM Leaderboard, 59:20
Well, 28 years old is not exactly correct, but we'll call it close enough. I like that, thank you very much, Llama 7B. Jeremy Howard, on his first local generation, 1:04:30
Something really important that I keep forgetting, everybody keeps forgetting, is that the instructions that are passed in, they actually always are in a particular format. And the format, believe it or not, changes quite a bit from fine tune to fine tune. Jeremy Howard, 1:08:15
If somebody asked me to send them a 100 word bio, that would actually probably be better than I would have written myself. Jeremy Howard, grading the retrieval augmented answer about himself, 1:13:25
So anyway, you can see the basic idea. This is called retrieval augmented generation, RAG. And it's a nifty approach, but you have to do it with some care. Jeremy Howard, right after it failed on a follow up question, 1:19:39
I've only tried this one, h2oGPT. I don't love it. It's all right. Jeremy Howard, 1:20:11
It also took me like an hour to figure out how to do it, and then an hour to actually do the training. And at the end of that we've actually got something which is converting prose into SQL based on a schema. Jeremy Howard, on the fine tune, 1:25:52
In some ways it's a frustrating time to get involved, because it's very early, and a lot of stuff has weird little edge cases and it's tricky to install. Jeremy Howard, closing, 1:30:03
Where this sits in the LLM Learning track
This is Part 4, "Ship something with it," and it is the hinge of the whole track. Everything before it explains what the model is and how it came to be: Karpathy's deep dive over the whole stack, 3Blue1Brown opening up attention one matrix at a time, Sasha Rush putting numbers on it, then the two builds, the tokenizer and GPT-2 reproduced, then the two talks on behavior, Schulman on RLHF and Olah on interpretability. From here the question changes from what it is to what you do with one.
Read in that order, this page pays off three earlier ones directly. Howard's three stage recipe is the pipeline Karpathy walks, named by the person who published it first, so Figure 1 is worth holding next to the Karpathy page. His token demonstration with tiktoken, where "They are splashing" becomes four tokens and the leading space belongs to the token, is the one minute version of the four hour tokenizer build. And his explanation for why GPT-4 answers confidently about things it cannot know, that the preference stage rewarded confidence and the annotators often could not spot a wrong answer, is the same mechanism Schulman spends an hour on, arriving from the user's side of the API rather than the researcher's.
It hands off forwards too. The sharpest thing Howard says about measurement is that the leaderboard everyone quotes is "a really fraught area," that its metrics are poorly aligned with real use, that leakage puts test questions into training sets, and therefore "you should always try things." He leaves that as a warning rather than a method, which is exactly the gap Hamel Husain and Emil Sedgh fill in Part 9 with assertions in CI, human review, and an aligned judge. And the one rung he does not climb here is autonomy: he builds a tool the model may call once, which is the first step of the ladder LangGraph walks all the way up in Part 10. Read 8, 9 and 10 together and you have the practical core of the track.
Resources mentioned
The talk's own materials
- fastai/lm-hackers, the repo for this talk, containing
lm-hackers.ipynb, the axolotl configsql.yml, and the tokenizing strategycontext_qa2.py - fast.ai and Jeremy Howard's site, plus his YouTube channel
- course.fast.ai, the free course whose first five lessons he names as the prerequisite
- The fast.ai Discord, and the
generativechannel he points people to
Papers
- Universal Language Model Fine-tuning for Text Classification, the ULMFiT paper, Howard and Sebastian Ruder, 2018
- GPT-4 Can't Reason, Konstantine Arkoudas, the paper whose examples he re runs
- The Flan Collection, which OpenOrca is built on
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- LoRA: Low-Rank Adaptation of Large Language Models and QLoRA: Efficient Finetuning of Quantized LLMs
- Orca and Platypus, the two lineages behind OpenOrca Platypus
- GSM8K and BIG-bench Hard, the benchmarks he names on the FastEval board
- WizardMath, one of his three model recommendations
Hosted models and the OpenAI API
- The OpenAI API docs, pricing, rate limits, and function calling
- nat.dev, Nat Friedman's playground with the show probabilities toggle, now open sourced as openplayground
- Bard, now Gemini, which pasted an image straight into the prompt and found its source
- Bing, which wrote his rate limit retry loop for free
Libraries
- tiktoken, transformers, datasets, accelerate, peft, PyTorch
- pydantic and the standard library's inspect, which together generate his function schemas
- fastcore, for
nested_idx - sentence-transformers and the Wikipedia-API package
- axolotl, the fine tuning harness
- llama.cpp and its Python wrapper, and MLC LLM
- h2oGPT, and its comparison of private GPT projects
Models and datasets on Hugging Face
- meta-llama/Llama-2-7b-hf and Llama-2-13b-hf
- TheBloke, and specifically Llama-2-7b-Chat-GPTQ, Llama-2-13B-GPTQ, OpenOrca-Platypus2-13B-GPTQ and Llama-2-7b-Chat-GGUF
- stabilityai/StableBeluga-7B from Stability AI, and StableBeluga2
- Open-Orca/OpenOrca-Platypus2-13B
- cognitivecomputations/dolphin-llama-13b
- BAAI/bge-small-en-v1.5, the 384 dimension embedding model behind the retrieval demo
- Open-Orca/OpenOrca, garage-bAInd/Open-Platypus and knowrohit07/know_sql
- The Open LLM Leaderboard and FastEval
Compute
- Free: Kaggle notebooks and Colab
- Rented: RunPod, Lambda Labs and vast.ai
People and things named in passing
- Alfred Hitchcock and The Birds, the Wikipedia article his 2017 model trained on
- The river crossing puzzle, and Steve Newman's variant of it
- Jeremy Howard's Wikipedia page and Tony Blair's, the two documents in the retrieval demo
An honest footnote
Four things worth saying once, at the end, for anyone working from this page rather than just reading it.
The caption track mangles names, and these are the ones that matter. The GPU rental service he recommends at 56:14 is vast.ai, not "fast.ai" as the captions have it, and that one is genuinely confusing because fast.ai is his own organization. "Discretization" throughout the local models section is quantization. "Fraud area" at 59:20 is "fraught area". Also corrected on this page, in his order: "openalker" is OpenOrca, "Sebastian Rooter" is Sebastian Ruder, "the bloke" is TheBloke, "Laura" is LoRA, "dolphin Lima 13B" is Dolphin Llama 13B, "lima.cpp" is llama.cpp, and "first.ai" and "faster AI" are both fast.ai. One slip is his own rather than the captions': he says "GTX 3090" and his notebook writes it the same way, but the card is an RTX 3090.
Two details come from the notebook rather than the talk. He says only "a Wikipedia python package" and "sentence transformer" out loud. The notebook shows the package is Wikipedia-API, imported as wikipediaapi, and the embedding model is BAAI/bge-small-en-v1.5, which is where the 384 dimensions come from. Likewise, the exact timings, token ids, cosine similarities and SQL output quoted on this page are the saved cell outputs in lm-hackers.ipynb, which is why they are precise rather than rounded.
One number does not match its source. He describes GPT-4 Can't Reason as covering 25 diverse reasoning problems. The paper's abstract says 21. This changes nothing about his argument, since he re ran the specific examples and showed the answers, but if you go and count them yourself, 21 is the number you will find.
The API he writes against has since moved, and the ideas have not. ChatCompletion.create, the functions parameter and the function_call response field are the pre 1.0 openai Python client. The current client is client.chat.completions.create, and functions became tools with tool_calls coming back, which allows more than one call at a time. The function message role is now tool. Everything structural in Figure 3 survives the rename: your code still generates a schema, the model still returns a request rather than an answer, your code still decides whether to run it, and the result still re enters as its own message rather than as an assistant turn. Two of his links have also drifted: axolotl moved from the OpenAccess AI Collective org, and the hosted nat.dev now redirects to Nat Friedman's personal site, with openplayground as the self hosted replacement.
None of that dents the talk. Its value was never the specific model names, and the parts that have aged best are the ones he derived rather than reported: that nothing in the training pipeline optimized for truth, that each token is another chance to compute so you should buy more of them before the answer, that a wrong turn in the context makes the next turn worse so you edit instead of arguing, that inference is bound by how fast weights move rather than how fast they multiply, and that a docstring is now a program. The 2023 prices are history. Those five are not.


