At a glance
Two speakers, one shipped product, and eighteen minutes on the single thing that separates a language model demo from a language model business. Emil Sedgh, CTO at Rechat, opens with the arc of Lucy, the assistant Rechat built into its real estate platform: a GPT 3.5 prototype wired up with the ReAct pattern that was slow, wrong most of the time, and "a majestic experience" when it worked, followed by a long stretch in which the team could change a prompt and have no idea whether they had improved anything or broken three other features. Hamel Husain, founder of Parlance Labs, takes the second half and walks through the recipe that got them out of it.
The recipe has a deliberate order, and the order is the argument. Start with unit tests and assertions written out of real observed failures, and run them wherever you already run tests. Log the results to whatever database or dashboard you already own. Log your traces, then actually look at them, which in Rechat's case meant building their own trace viewing and labeling app because the off the shelf tools carried too much friction. Bootstrap your test cases by having a model roleplay as your Customer. Exercise the whole loop with plain prompt engineering so you find out whether the loop works. Only then reach for an LLM as a judge, and only after you have measured its verdicts against a human's.
Husain is blunt about where teams go wrong: they do not look at their data, they talk about tools before they have a process, they reach for off the shelf conciseness and toxicity scores instead of writing evals for their own domain, and they jump to an LLM judge while there are still plenty of cheap assertions sitting unwritten in the logs. Sedgh closes with the payoff. The success rate went up fast, the three hardest behaviors in the product turned out to need fine tuning rather than clever prompting, and the marketing workflow he demonstrates on stage, which he estimates takes a non technical agent a couple of hours by hand, runs in about a minute.
It is the least glamorous talk in the LLM Learning track and the one that decides whether anything else in the track ships.
The talk, rebuilt
0:00 Rechat, and the brilliant idea
Sedgh introduces himself as CTO at Rechat and frames the talk in four parts: the product they built, the challenges they faced, how their eval framework came to the rescue, and the results.
The setup matters because it explains why nothing off the shelf could have measured this product. Rechat's application is built for real estate agents and brokers, and by the time they started thinking about AI it already carried a wide feature surface: contact management, email marketing, social marketing, "whatever," as Sedgh puts it. Two things fell out of having built all of that. They had a lot of internal APIs, and they had a lot of data.
So, at 0:30, the decision, delivered with a straight face: "naturally we came to the unique and brilliant idea that we need to build an AI agent for our real estate agents." The joke is doing real work. Every company with internal APIs and a data warehouse reached the same conclusion in the same quarter, which is exactly why the rest of the talk is about the part nobody copied.
1:03 The prototype: slow, wrong, and majestic
Sedgh rewinds a year, to 2023. They built a prototype on the original GPT 3.5 and the ReAct framework, the reasoning and acting loop in which the model alternates between thinking out loud and calling a tool. (The captions render this as "react framework," which reads like the JavaScript library. It is ReAct, the agent pattern.)
His verdict on that prototype is the most quoted line in the first half: "it was very very slow and it was making mistakes all the time, but when it worked it was a majestic experience, it was beautiful experience." That sentence is the trap the whole talk is about. A system that is occasionally magnificent gives you no information about how often it is magnificent, and the feeling of the good runs is strong enough to drown out the arithmetic.
They concluded they had a product in a demo state and now had to take it to production, and that is when they started working with Husain.
Sedgh shows the basic shapes of what agents ask Lucy to do:
- Create a contact for me with this information.
- Send an email to somebody, with some instructions.
- Find me some listings, "because that's what real estate agents tend to do."
- Create a website for me.
Four categories, and already four different notions of correct. A created contact is checkable against a database. An email is a rendering problem and a tone problem at once. A listing search has a right answer only relative to the criteria. A generated website has almost no single right answer at all. No public leaderboard has ever scored any of this, which is the whole case for domain specific evaluation.
2:01 The improvement phase, where the floor fell out
Then comes what Sedgh calls the improvement of language model phase, and the two problems that defined it.
The first problem was that they could not tell whether changes helped. "The problem was when we tried to make changes to see if we can improve it, we didn't really know if we're improving things or not. We would make a change, we would invoke it a couple of times, we would get a feeling that yeah it worked a couple of times, but we didn't really know what the success rate or failure rate was. Is it going to work 50% of times or 80% of times?" He follows it with the business consequence, which is the sentence an executive would care about: "it's very difficult to launch a production app when you don't really know how well it's going to function."
Note what the 50% versus 80% figure is and is not. It is a rhetorical pair, not a measurement. The point is precisely that they had no measurement, and the gap between those two numbers is the gap between a product and an embarrassment.
The second problem was regression. Even when a change felt like an improvement, "the moment we changed the prompts it was likely that it's going to break other use cases." A prompt is a single shared surface that every feature reads from, so a fix aimed at one behavior perturbs all the others, and without a test suite nobody finds out until a Customer does.
His summary of the state they were in, at 2:31: "we were essentially in the dark." He hands over to Husain.
3:01 Vibe checks get you to one and no further
Husain starts by naming what Sedgh had actually been doing, and gives it credit. "What Emil described is he was able to use prompt engineering, implement RAG, agents, so on and so forth, and iterate with just vibe checks really fast to go from zero to one, and this is a really common approach to building an MVP. It actually works really well for building an MVP."
Then the turn: "however, in reality this approach doesn't work for that long at all. It leads to stagnation, and if you don't have a way of measuring progress you can't really build."
The word to notice is stagnation, not failure. The vibe check loop does not blow up. It plateaus. You keep working, you keep shipping prompt changes, and the product stops getting better in any way you can demonstrate, because every change you make is as likely to be a wash or a regression as an improvement and you have no instrument that can tell the three apart.
He sets out the talk's agenda: a systematic approach you can use to improve your AI consistently, how to avoid the common traps, and resources for learning more, "because you can't learn everything in a 15-minute talk." Then he puts up the recipe diagram and immediately defuses it: "you don't have to fixate too much on the details of this diagram because I'm going to be walking through it slowly."
That diagram is the spine of everything that follows, so here it is rebuilt from what he says as he walks it.
4:18 Level one: unit tests and assertions
The first level is deliberately boring, and Husain spends real time on why boring is the point.
"A lot of people are familiar with unit tests and assertions if you have been building software, but for whatever reason people tend to skip this step, and it's kind of the foundation for evaluation systems." Then the instruction, which is the most actionable sentence in the talk: "you don't want to jump straight to LLM as a judge or generic evals, you want to try to write down as many assertions and unit tests as you can about the failure modes that you're experiencing with your large language model, and it really comes from looking at data."
That last clause is the mechanism. You do not sit down and imagine failure modes. You read traces, you find a bad output, and you turn that specific bad output into a check that can never pass silently again.
The slide shows real assertions Rechat wrote, "based upon failure modes that we observed in the data," and he is explicit that these are a sample: "these are not all of them, there's many of these, but these are just examples." The ones he reads out:
- Testing that the agents are working properly, "so emails not being sent."
- Invalid placeholders in generated output.
- "Other details being repeated when they shouldn't."
Then he waves off the specifics on purpose: "the details of these specific assertions don't matter. What I'm trying to drive home is this is a very simple thing that people skip but it's absolutely essential, because running these assertions give you immediate feedback and are almost free to run, and it's really critical to your overall evaluation system if you can have them."
Two properties do the work there, and they are properties no other level has. Immediate feedback, and almost free. Those two together are what let a check run on every commit, which is what makes it a ratchet rather than a report.
5:33 Where to run them, and where the results go
"How do you run the assertions? One very reasonable way is to use CI." Then, immediately, the honest caveat: "you can outgrow CI and it may not work as you mature, but one theme I want to get across is use what you have when you begin, don't jump straight into tools."
The second half of the same instruction is about storage. "Another thing that you want to do with these assertions and unit tests is log the results to a database, but when you're starting out you want to keep it simple and stupid, use your existing tools."
Rechat's case is the whole argument in one sentence: they were already using Metabase, so the assertion results went to Metabase, and they used Metabase dashboards to visualize and track results, "so that we could see if we're making progress on these dumb failure modes over time."
The phrase "dumb failure modes" is not dismissive, it is a category. These are the failures that are embarrassing rather than interesting: a missing disclaimer, a leaked UUID, a placeholder that never got filled. They are also the failures that generate support tickets, and they are the only class of failure you can regression test for free.
He repeats the rule, because it is the one he expects the audience to break: "my recommendation is don't buy stuff, use what you have when you're beginning, and then get into tools later."
6:34 Logging traces, and the one tool he says buy on day one
The next piece is logging and human review, and here Husain makes his single exception to the do not buy anything rule.
"It's important to log your traces. There's a lot of tools that you can use to do this. This is one area where I actually do suggest using a tool right off the bat." The slide lists commercial and open source options. Rechat's choice: LangSmith, whose observability model is built around exactly this unit, the trace that records a run end to end rather than a single model call.
The exception is worth understanding rather than just noting. Tracing is plumbing with a well defined output and no domain judgment in it, and writing your own buys you nothing. Everything above tracing involves a verdict about your own product, and that is where he says build.
Which brings the hinge of the talk: "more importantly, it's not enough to just log your traces, you have to look at them, otherwise there's no point in logging them."
7:05 Build your own data viewer
"One nuance here is that looking at your data is so important that I actually recommend building your own data viewing and annotation tools in a lot of cases, and the reason is because your data and application are often very unique. There's a lot of domain specific stuff in your traces."
Rechat's experience is given plainly: "in Rechat's case we found that tools had too much friction for us, so we built our own kind of little application." On how: "you can do this very easily in something like Gradio, Streamlit. I use Shiny for Python. It really doesn't matter."
That "it really doesn't matter" is load bearing. The choice of framework is not the decision. The decision is that the tool is yours and therefore can hold things no vendor would build:
- Filters that are specific to Rechat, cutting the data in ways only someone inside the product would want.
- All the Rechat specific metadata attached to each trace, on the same screen. His reason: "where I don't have to hunt for information to evaluate a trace."
- Labeling, not just viewing. "This is not only a kind of a data viewing app, this is also a data labeling app where it facilitates human review."
Then, at 8:07, the line he tells the audience to keep if they keep nothing else:
"This is the most important part. If you remember anything from this talk, it is you need to look at your data, and you need to fight as hard as you can to remove all friction in looking at your data, even down to creating your own data viewing apps if you have to."
And the failure mode if you do not, which is a claim about people rather than about software: "it's absolutely critical. If you have any friction in looking at data, people are not going to do it, and it will destroy the whole process, and none of this is going to work."
That is the real argument for a custom tool, and it is not a technical argument. A review tool that requires fifteen clicks to reach a trace does not get used a little less. It gets used zero times, and every level above it is built on its output.
8:38 Synthetic data: an LLM roleplaying your Customer
Husain anticipates the obvious objection. You have unit tests, logging and human review, "and you might be wondering, okay, you have these tests, what about the test cases? What do we do about that, especially when you're starting out, you might not have any users?"
The answer: "you can use LLMs to synthetically generate inputs to your system."
Rechat's implementation, verbatim: "in Rechat's case we basically use an LLM to roleplay as a real estate agent and ask questions as inputs into Lucy, which is their AI assistant, for all the different features and the scenarios and the tools, to get really good test coverage."
Three axes are named in that sentence, and they are the whole method: features, scenarios, and tools. You do not ask a model for test cases in general. You enumerate the surface of your product, then generate inputs per cell. He frames it modestly as a bootstrap: "using LLMs to synthetically generate inputs is a good way to bootstrap these test cases."
9:40 Test the eval system by trying to improve the product
At this point the minimum viable eval system exists: unit tests, logged traces, human review. Husain calls it exactly that, twice: "when you have a very minimal setup like this, this is like the very minimal thing, like a very minimal evaluation system, like bare bones."
What he says to do next is the step almost nobody writes down. Do not go build more eval machinery. Go use the machinery you just built, on purpose, to find out whether it works.
"What you want to do when you first construct that is you want to test out the evaluation system, so you want to do something to make progress on your AI, and the easiest way to try to make progress on your AI is to do prompt engineering. So what you should do is go through this loop as many times as possible."
The checklist he gives for what that exercise is actually testing:
- Is your test coverage good?
- Are you logging your traces correctly?
- "Did you remove as much friction as possible from looking at your data?"
"This will help you debug that, but also give you the satisfaction of making progress on your AI as well." Prompt engineering is the cheapest possible payload to push through a new pipeline, which makes it the right smoke test for the pipeline itself, and it also pays the team back immediately so the pipeline does not feel like overhead.
10:42 The superpower you get for almost free
Then a widening of the frame, at 10:42: "one thing I want to point out is the upshot of having an evaluation system is you get other superpowers for almost free."
The superpower is fine tuning, and the argument is about where the labor in fine tuning actually sits: "all of the work in fine-tuning, or most of the work, is data curation."
If that is true, then an eval system is a fine tuning data factory that happens to also tell you your success rate. He describes both halves of the flow:
- Good cases: "you can use your eval framework to kind of filter out good cases and feed that into your human review, like we showed with that application, and you can start to curate data for fine tuning."
- Failed cases: "for the failed cases you have this workflow that you can use to work through those and continuously update your fine-tuning data."
Note that the failed cases are not discarded. The annotation app is where a human can take a bad output and turn it into the output that should have been produced, which converts a failure into a training example rather than just a bug report.
And then the compounding claim, which is the economic case for the whole system: "what we've seen over time is that the more comprehensive your eval framework is, the cost of human review goes down, because you're automating more and more of these things and getting more confidence in your data."
Read that carefully, because it runs against the intuition that evals are a growing tax. The tax falls over time. Every assertion you write is a class of trace a human never has to look at again.
11:37 LLM as a judge, and the alignment discipline
"Once you have this setup, now you're in a position to know whether or not you're making progress or not. You have a workflow that you can use to quickly make improvements and you can start getting rid of those dumb failure modes. But also now you're set up to move into more advanced things like LLM as a judge, because you can't express everything as an assertion or a unit test."
That clause is the entire justification for a judge, and it is a narrow one. The judge exists to cover the residue: the qualities that no string check, schema validation or database lookup can express. Was the email any good. Did it sound like this brokerage. Was declining to act the right call.
He declares the topic out of scope and then says the one thing he refuses to leave out: "LLM as a judge is a deep topic, just outside the scope of this talk, but one thing I want to point out is it's very very important to align the LLM judge to a human, because you need to know whether you can trust the LLM as a judge. You need a principled way of reasoning about how reliable the LLM as a judge is."
The procedure, deliberately unimpressive: "what I like to do is again keep it simple and stupid. I like to use a spreadsheet often, don't make it complicated. But what I do is have a domain expert label data, label the critique and critique data, and keep iterating on that until my LLM as a judge is in alignment with my human judge, and I have high confidence that the LLM judge is doing what it's supposed to do."
Three things are specified there and they are all worth separating out. The labeler is a domain expert, not an engineer. The label is not only a verdict, it is a critique, so the judge has the reasoning to imitate and not just the score. And the artifact that holds the comparison is a spreadsheet, because the comparison is a handful of columns and the moment it becomes a platform it stops getting done.
What the talk does not give is a number. It states no agreement rate, no sample size, no threshold for "in alignment." Husain's own written treatment of the same method does, and the scaffolding section further down this page quotes those figures with their source, clearly marked as coming from the writing rather than from the stage.
12:45 Four ways people get this wrong
"I'm going to go through some common mistakes that people make when building LLM evaluation systems."
One, not looking at your data. "It's easier said than done, but people don't do the best job of doing this, and one key to unlocking this is to remove all the friction, as I mentioned before." Same diagnosis, same fix, stated twice in one talk, which tells you how much weight he puts on it.
Two, focusing on tools rather than processes. He rates this "just as important" as the first. "If you're having a conversation about evals and the first thing you start thinking about is tools, that's a smell that you're not going to be successful in your evaluations. People like to jump straight to the tools: tell me about the tools, what tools should I use."
The reasoning is sharper than the usual anti vendor complaint, because it is about your ability to evaluate the vendor: "it's really important to try not to use tools to begin with, and try to do some of these things manually with what you already have, because if you don't do that you won't be able to evaluate the tools, and you have to know what the process is before you jump straight into the tools, otherwise you're going to be blindsided."
A team that has never labeled a trace by hand has no basis on which to compare two labeling products. They will buy on the demo and discover the mismatch after the integration.
Three, generic off the shelf evals. "People using generic evals off the shelf. You don't want to reach for generic evals, you want to write evals that are very specific to your domain. Things like conciseness score, toxicity score, all these different evals you can get off the shelf with tools, you don't want to go directly to those. That's also a [smell] that you are not doing things correctly."
He then declines the easy overstatement: "it's not that they're not valuable at all, it's just that you shouldn't rely on them because they can become a crutch." A conciseness score is a real number. It is just not a number that tells you whether Lucy built the right marketing email for the right listing under the right brand.
Four, reaching for an LLM judge too early. "The other common mistake is with LLM as a judge, and using that too early. I often find that if I'm looking at the data closely enough I can always find plenty of assertions and failure modes. It's not always the case, but it's often the case. So don't go to LLM as a judge too early, and also make sure you align LLM as a judge with a human."
Note the hedge he puts on his own claim. "It's not always the case, but it's often the case" is him telling you this is a strong prior from consulting experience rather than a law. The practical form of the prior: if you cannot write assertions, it is more likely that you have not read enough traces than that your problem is irreducibly subjective.
He hands back to Sedgh for the results.
| Grader | What it can judge | Cost and speed | What it cannot tell you | When the talk says to reach for it |
|---|---|---|---|---|
| Assertion or unit test | Anything expressible in code: schema validity, a placeholder that is still a placeholder, a UUID that leaked into prose, whether the email actually got sent, a detail repeated that should appear once | "Almost free to run," immediate feedback, runs in CI on every commit | Whether the output was any good. It can confirm an email is well formed and say nothing about whether it was the right email | FIRST, ALWAYS The foundation. Write as many as you can, sourced from failures you observed in real traces |
| Human review in a custom app | Everything, including the things nobody has thought to specify yet. The only grader that can discover a new failure mode rather than detect a known one | The expensive one. Bounded by how many traces a domain expert will actually sit through, which is bounded by friction | Nothing, but it does not scale, and it is the one step that quietly stops happening if the tool is annoying | SECOND, AND NEVER SKIPPED "If you remember anything from this talk." It is also the input every later level depends on |
| LLM as a judge | The residue: the qualities no assertion can express. Used because "you can't express everything as an assertion or a unit test" | Cheap enough to run over far more traces than a human can read, with a real per call cost and a prompt to maintain | Whether it is right, until you have measured its verdicts against a human's on labeled data | LAST, AND NOT EARLY Only after the assertions are written, and only once aligned to a domain expert's labels. A deep topic he puts out of scope |
| Generic off the shelf eval | Domain independent properties: conciseness, toxicity, and the rest of the standard battery | Cheapest to adopt, since adoption is an import statement | Anything about your product. No public metric scores "did it build the right marketing email for this listing under this brand" | NOT AS A STARTING POINT "Not that they're not valuable at all," but relying on them "can become a crutch" and reaching for them first is a smell |
15:11 What it bought them
Sedgh takes the last three minutes with the results, and he is careful about what he claims.
"After we got to the virtuous cycle that Hamel just displayed, we managed to rapidly increase the success rate of the LLM application. Without the eval framework, a project similar to this seemed completely impossible for us."
Two separate statements there, and only one of them is a measurement claim. The success rate went up rapidly, with no number attached. The second is stronger and stranger: not that the eval framework made the project faster or cheaper, but that without it the project looked impossible. Given the first half of the talk, that is coherent. They were in the dark, every change was a coin flip, and a product you cannot measure is a product you cannot responsibly launch.
15:34 Why few shot prompting was not enough
Then Sedgh pushes back on a claim he says he keeps hearing: "one thing that I've started to hear a lot is that few shot prompting is going to replace fine tuning, or some notions like that. In our case, we never managed to get everything that we wanted by few shot prompting, even using the newer and smarter [models]."
He frames it as something he would have preferred to be true, and the aside is the best line in the second half: "I wish we could. I've seen a lot of judgment of companies and products being just ChatGPT wrappers. I wish we could just be a ChatGPT wrapper and manage to extract the experience we want for our users, but we never had that opportunity, because we had some really difficult cases."
The three difficult cases are the substance of this section, and each one is a specific product requirement rather than a benchmark.
One, mixing natural language with user interface elements in a single output. "One of the things that we wanted our agent to be able to do was to mix natural language with user interface elements like this inside the output, and this essentially required us to mix structured output and unstructured output together. We never managed to get this working without fine tuning reliably."
This is the hardest of the three and the least discussed in the wider literature. A response that is partly prose and partly a rendered widget is not a JSON schema problem and not a free text problem, it is both interleaved, and a model that is good at either one separately will drift at the seams.
Two, asking the user for more input. "Sometimes the user asks, in a case like this, do this for me, but the agent can just [not] do that, it needs some sort of feedback, more input from the user. Again, something like this was very difficult for us to execute on, especially given the previous challenge of injecting user interfaces inside the conversation."
Knowing when not to act is a behavior, not a capability, and it is the behavior most likely to be trained out of a model by instruction tuning that rewards helpfulness. It is also compounded by the first problem: the right way to ask for missing input is usually a form, which means asking well requires emitting interface mid conversation.
Three, complex multi tool commands. "The third reason that we had to fine tune was complex commands like this." The requirement: take an input that "requires using like five or six different tools to be done," break it down into many different function calls, and execute them.
16:59 The complex command, start to finish
Sedgh plays a short video of one such command running, and narrates it. This is the most concrete thing in the talk, so it is worth writing out as the sequence it is.
The request, in his words: find me some listings with some criteria, then create a website, "that's what real estate agents sometimes do for their listings that they're responsible for," and also an Instagram post so they can market it, and do that only for the most expensive listing of the three.
What the application did:
- Found three listings matching the criteria.
- Selected the most expensive of the three, which is the conditional buried in the middle of the sentence.
- Created a website for that listing.
- Created and rendered an Instagram post video for it.
- Prepared an email to Husain containing all the information about the listings, the website that was created, and the Instagram story that was created.
- Invited Husain to a dinner.
- Created a follow up task.
Seven actions from one sentence, with a comparative filter applied partway through and three generated artifacts referenced correctly in a fourth artifact. Notice how much of the difficulty is in step 5, where the email has to contain working references to things that did not exist when the request was made.
His estimate of the value, at 17:57: "creating something like this for a non savvy real estate agent may take a couple of hours to do, but using the agent they can do it in a minute."
That is an informal estimate from the stage rather than a measured benchmark, and he does not present it as one. It is also the only pair of concrete figures in the talk, so it is worth looking at directly.
Sedgh closes by connecting that back to the talk's thesis: "that essentially was not going to be possible without us using a comprehensive eval framework." Then, with an eye on the clock, Husain's verdict on the timing: "nailed the timing, thank you guys."
How Rechat got from prototype to this
The talk tells the story in two voices and out of order, so here is the arc on one axis. Everything dated 2023 or 2024 below is from the talk itself unless the entry says otherwise.
- 2023Lucy starts as something much smaller than an agent: a tool for answering questions about transaction forms, according to the AI Engineer speaker page for Sedgh. Rechat introduced Lucy publicly in September 2023.
- 2023The prototype. Built on the original GPT 3.5 with the ReAct pattern, on top of the internal APIs and data Rechat already had. "Very very slow," "making mistakes all the time," and "a majestic experience" when it worked. Demo state reached.
- 2023The improvement phase stalls. Changes cannot be told apart from noise. No known success rate. "Is it going to work 50% of times or 80% of times?" Every prompt change risks breaking other use cases. "We were essentially in the dark."
- 2023Rechat brings in Husain to get the thing production ready. The vibe check loop had taken them from zero to one and then flattened out.
- Level 1Assertions written from observed failures, run in CI, results logged into the Metabase instance they already had, and tracked over time so progress on the dumb failure modes is visible.
- Level 2Traces logged to LangSmith, then the custom trace viewing and labeling app in Shiny for Python, with Rechat specific filters and metadata on one screen so a reviewer never hunts for context.
- Level 2Synthetic inputs from an LLM roleplaying a real estate agent, generated across features, scenarios and tools, to get coverage before and beyond what real traffic provides.
- LoopPrompt engineering as the smoke test for the eval system itself, run as many times as possible, checking coverage, logging and friction while also making the product better.
- Level 3LLM as a judge, for the residue no assertion can express, with a domain expert labeling verdicts and critiques in a spreadsheet until the judge agrees with the human.
- PayoffFine tuning becomes available, because the eval system is also a data curation pipeline. Good cases filtered and kept, failed cases corrected in the annotation app and fed back. Human review cost falls as coverage grows.
- Jun 2024The talk, at AI Engineer World's Fair 2024 in San Francisco, June 25 to 27. Success rate "rapidly" increased, three behaviors shipped that few shot prompting never delivered, and the seven action marketing command running live on stage.
The scaffolding: what it takes to build this
The talk is eighteen minutes and Husain says outright that you cannot learn everything in fifteen. What follows is the operational version of his recipe, assembled from what the talk specifies plus the figures from the essay the same authors published on the same system. Every number below comes from Your AI Product Needs Evals (Husain, 29 March 2024), not from the stage, and is marked as such. The talk states no counts, no sample sizes and no agreement rates.
Level one, concretely
Start from a failure you actually saw, not from a quality dimension you imagined. The essay's Rechat examples show what that produces, and they are more mundane than most eval discussions:
- Scenario coverage on a feature. For the listing finder, three assertions on the shape of the result: exactly one match (
len(listing_array) == 1), several matches (len(listing_array) > 1), and no matches (len(listing_array) == 0). Three branches of one feature, three checks, no model involved. - A generic structural check. A regular expression asserting that no UUID ever appears in text shown to a user. One line of code that closes an entire class of leak.
- The talk adds the three it shows on the slide: emails that were supposed to be sent and were not, invalid placeholders, and details repeated that should appear once.
How many is enough? The essay says Rechat has "hundreds of these unit tests." The talk only says "as many as you can," and that the slide examples are a sample of many more.
A practical consequence of the shape of these tests: they do not need a frontier model, a scoring rubric or a vendor. They need a parser, a database handle and a regular expression, which is why they can run on every commit.
Where the test cases come from
The talk gives the three axes and the method. The essay gives the prompt shape, and it is worth seeing how low tech it is: ask a model for 50 instructions a real estate agent could give an assistant to create a CRM contact, covering the fields the product actually has (name, phone, email, partner name, birthday, tags, company, address, job), and for each one also generate a second instruction that looks that contact back up. Return a JSON array of the pairs.
Three things make that work, and all three are generalizable:
- The enumeration is yours, not the model's. Features, scenarios and tools come out of your product, and the model only fills in phrasings. Asking a model to invent test categories gets you the categories everyone has.
- The field list is the coverage. Naming the nine contact fields is what forces cases that exercise the awkward ones. Nobody writes a birthday test by hand.
- The pairs make the case checkable. Generating the lookup instruction alongside the creation instruction gives you an assertion for free, because the second instruction either finds what the first created or it does not.
The annotation app, as a spec
The talk says build one and names three frameworks. The essay describes what Rechat's actually contains, which converts "build your own tool" into a requirements list:
- The feature and scenario being exercised, shown on the trace.
- Whether the input was synthetic or came from a real user.
- Filters over all of that.
- Links straight out to the CRM record and to the trace log, so verifying a claim does not mean switching products and searching.
- An editable final output, so a reviewer can correct a bad response in place and that correction becomes fine tuning data.
- A binary good or bad label rather than a numeric score. Start there.
The essay's claim on effort is that a tool like this can be built in under a day with Gradio or Streamlit. The binary label detail is the one most teams get wrong: a 1 to 5 scale invites reviewers to park everything on 3, and the disagreements between two reviewers on a 5 point scale are usually about the scale rather than about the output.
Aligning the judge, with the numbers the talk leaves out
The talk specifies the method: a domain expert labels verdicts and critiques, you iterate the judge prompt in a spreadsheet until it agrees with the human. The essay supplies the operating parameters:
- Batch size for human grading: 25 to 50 examples at a time. For each one the grader writes a critique, an outcome, and the response that should have been produced.
- The judge produces the same two artifacts the human does, a written critique and a binary good or bad outcome, so the comparison is like for like.
- Measure agreement, and know when raw agreement lies. The essay reports raw agreement because its dataset was roughly balanced at about 50% failures. It explicitly warns that with imbalanced classes you should measure precision and recall, or the true positive and true negative rates, separately. A judge that calls everything good scores 95% raw agreement on a dataset that fails 5% of the time, and is worth nothing.
- For a reference point on what good looks like, Husain's full guide to LLM as a judge reports better than 90% agreement between judge and expert in its worked example. That post was published in October 2024, after this talk, and its example is a different product.
The discipline the talk insists on, and the one easiest to let slide, is that the judge's credibility is borrowed and expires. It is trusted exactly as far as its last measured agreement with a human, on data that resembles what you are running it on now.
Where it sits in the development loop
Pulling together what the talk specifies about the mechanics:
- Assertions run in CI, on every change, because they are almost free and the feedback is immediate. Husain's caveat is that CI may not be where you run them forever, and that is fine: use what you have.
- Assertion results go to a time series you can look at, which for Rechat meant the Metabase they already owned, because the question is not only "did this commit pass" but "are we getting better at this failure mode over time."
- Traces are logged with a tool you buy, from day one, and LangSmith was the choice here.
- Human review happens in a tool you build, continuously, because it is both the discovery mechanism for new failure modes and the ground truth for everything automated.
- The judge runs over the traces a human will never read, and gets re measured against fresh human labels as the product moves.
One thing the talk does not cover: a third level in the written version of this material. The essay's Level 3 is A/B testing, checking that the product actually drives the behavior you want with real users, and it says that is usually for a more mature product and can wait. On stage, the advanced level is the LLM judge and A/B testing does not come up. If you have read the essay and the talk and found the levels numbered differently, that is why.
| A public benchmark or off the shelf eval | A domain specific eval set, built the way this talk describes | |
|---|---|---|
| What the score means | Performance on a fixed, shared task: a conciseness score, a toxicity score, a leaderboard rank | Whether your product did the thing your Customer asked, on the features, scenarios and tools your product actually has |
| Where the cases come from | Someone else's dataset, built before your product existed and with no knowledge of it | Failures you observed in your own traces, plus synthetic inputs enumerated over your own feature surface |
| Who decides correct | The benchmark author, once | A domain expert on your side, continuously, in an app built so that reviewing is frictionless |
| What it catches | Gross regressions and broad model level differences | THE ACTUAL BUGS Unsent emails, unfilled placeholders, repeated details, leaked UUIDs, a listing search that returns three results when it should return one |
| What it misses | EVERYTHING THAT MATTERS TO YOU No public metric scores "did it build the right marketing email for this listing under this brand," because no public metric knows your brand, your listings or your CRM | Anything you have not thought of yet, which is exactly why human review of real traces never stops being the input |
| Cost profile | Near zero to adopt, which is the whole appeal and the whole hazard | Real work up front, falling over time as coverage grows and more of the review automates |
| The talk's verdict | "Not that they're not valuable at all," but reaching for them first is a smell, and relying on them "can become a crutch" | The title of the talk. Write evals that are very specific to your domain |
Where it stands, and what is opinion
Separating the three kinds of claim in the talk, because they carry different weight.
Results, stated without numbers. Sedgh says the success rate increased rapidly, that a project like this seemed impossible without the eval framework, and that the three hard behaviors were unreachable with few shot prompting on the models available to them. None of these come with a figure, a baseline or a date, and the talk does not pretend otherwise. The only quantities on stage are the "couple of hours" versus "a minute" estimate for the marketing workflow, which Sedgh offers as an observation about his Customers rather than a benchmark, and the rhetorical "50% or 80%" that exists precisely because they had no measurement.
Engineering advice backed by a real deployment. The order of the levels, the assertions in CI, the results tracked in a tool you already own, the bought tracing tool, the built annotation tool, the synthetic inputs across features and scenarios and tools, the judge aligned against a domain expert. These are reports of what a shipping product does, and they hold together with the written account the same authors published independently.
Opinion, held strongly. Several of the sharpest lines are judgments from consulting experience, not findings, and Husain generally marks them as such:
- That talking about tools first is "a smell that you're not going to be successful." A heuristic about teams, offered as one.
- That you should build your own data viewing app "in a lot of cases." Hedged, and the hedge matters, because a product with less domain structure in its traces gets less from a custom viewer.
- That there are almost always more assertions available if you look at the data closely enough. He explicitly qualifies this: "it's not always the case, but it's often the case."
- That generic evals become a crutch. A claim about incentives rather than about accuracy.
- Sedgh's rejection of few shot prompting as a replacement for fine tuning is about his three specific requirements on the models of mid 2024, not a general claim, and he says he wishes it had gone the other way.
What the talk does not cover, and says so. LLM as a judge in any depth, declared out of scope. A/B testing, which the written version makes its third level, does not come up at all. And there are no figures for the eval set itself: no case count, no labeler count, no agreement rate, no cost. The scaffolding section above fills those from the essay, with the source marked each time, because filling them from imagination would be the exact failure this talk is about.
On the captions. The automatic transcript mangles most of the proper nouns in this talk, so for the record: "haml" and one "Hammer" are Hamel Husain, "Emil s" is Emil Sedgh, "reat" and "rat" and "rehat" are Rechat, "lsmith" is LangSmith, "streamlet" is Streamlit, "cplay" is roleplay, "F shot" is few shot, "ChatGPT rapper" is ChatGPT wrapper, and "react framework" is the ReAct agent pattern rather than the JavaScript library. Quotations above are verbatim except for restored punctuation and capitalization, with any word corrected or supplied shown in square brackets.
Key takeaways
- The failure is not that the vibe check loop breaks, it is that it plateaus. You keep shipping prompt changes and the product stops improving in any way you can demonstrate, because you cannot tell an improvement from a regression from noise.
- Write assertions and unit tests first, and source them from failures you actually observed in traces rather than from imagined quality dimensions. They are almost free, the feedback is immediate, and that combination is what lets them gate every commit.
- Use what you already have when you start. Run the assertions in CI, log their results to the dashboard tool you already own, and do not buy an eval platform before you have a process to evaluate it against.
- Buy the tracing tool on day one. Build the review tool yourself. Tracing is plumbing with no domain judgment in it; reviewing is nothing but domain judgment.
- "If you remember anything from this talk, it is you need to look at your data." The corollary is the operational one: any friction in looking at data means people will not do it, and every level above human review is built on its output.
- Bootstrap test cases by having a model roleplay as your Customer, enumerated across your features, your scenarios and your tools. You supply the enumeration, the model supplies the phrasings.
- Once the minimal system exists, test it by doing plain prompt engineering through it as many times as possible. The cheapest payload is the right smoke test for a new pipeline, and it pays the team back immediately.
- An LLM judge is for the residue that no assertion can express, and it is worth exactly as much as its last measured agreement with a human expert. Align it in a spreadsheet; do not build a platform for a comparison that is four columns wide.
- The four common mistakes: not looking at your data, talking about tools before you have a process, reaching for generic off the shelf scores, and going to an LLM judge while cheap assertions are still sitting unwritten in your logs.
- The eval system pays for itself twice. Human review cost falls as automated coverage grows, and the curated data it produces as a byproduct is what makes fine tuning possible at all, since most of the work in fine tuning is data curation.
- Rechat's three hardest requirements all needed fine tuning rather than prompting: interleaving rendered interface with prose in one response, knowing when to stop and ask the user for more input, and decomposing one sentence into five or six tool calls.
Chapters
- 0:00:00 Introduction to Rechat
- 0:01:03 Initial prototype challenges
- 0:02:01 The need for evaluation
- 0:03:01 Systematic approach to evals
- 0:04:18 Unit tests and assertions
- 0:05:55 Logging and data usage
- 0:06:40 Human review and tooling
- 0:08:17 Synthetic data generation
- 0:09:37 Prompt engineering loop
- 0:11:37 Advanced evaluation
- 0:12:57 Common evaluation pitfalls
- 0:15:11 Results of the framework
- 0:15:34 Importance of fine-tuning
- 0:16:59 Complex task automation
Notable quotes
- "Naturally we came to the unique and brilliant idea that we need to build an AI agent for our real estate agents." Emil Sedgh, 0:30, on how Rechat's internal APIs and data led where everyone else's did.
- "It was very very slow and it was making mistakes all the time, but when it worked it was a majestic experience, it was beautiful experience." Sedgh, 1:03, on the GPT 3.5 prototype.
- "We would make a change, we would invoke it a couple of times, we would get a feeling that yeah it worked a couple of times, but we didn't really know what the success rate or failure rate was. Is it going to work 50% of times or 80% of times?" Sedgh, 2:01.
- "We were essentially in the dark." Sedgh, 2:31.
- "This approach doesn't work for that long at all. It leads to stagnation, and if you don't have a way of measuring progress you can't really build." Hamel Husain, 3:01, on the vibe check loop that gets an MVP out the door.
- "You don't want to jump straight to LLM as a judge or generic evals. You want to try to write down as many assertions and unit tests as you can about the failure modes that you're experiencing with your large language model, and it really comes from looking at data." Husain, 4:33.
- "Running these assertions give you immediate feedback and are almost free to run." Husain, 5:03.
- "Use what you have when you begin, don't jump straight into tools." Husain, 5:33.
- "Don't buy stuff, use what you have when you're beginning, and then get into tools later." Husain, 6:03.
- "It's not enough to just log your traces, you have to look at them, otherwise there's no point in logging them." Husain, 7:05.
- "In Rechat's case we found that tools had too much friction for us, so we built our own kind of little application." Husain, 7:36, on why a custom trace viewer beat the off the shelf options.
- "This is the most important part. If you remember anything from this talk, it is you need to look at your data, and you need to fight as hard as you can to remove all friction in looking at your data, even down to creating your own data viewing apps if you have to." Husain, 8:07.
- "If you have any friction in looking at data, people are not going to do it, and it will destroy the whole process, and none of this is going to work." Husain, 8:38.
- "In Rechat's case we basically use an LLM to roleplay as a real estate agent and ask questions as inputs into Lucy, which is their AI assistant, for all the different features and the scenarios and the tools, to get really good test coverage." Husain, 9:08.
- "The upshot of having an evaluation system is you get other superpowers for almost free." Husain, 10:42, introducing fine tuning.
- "All of the work in fine-tuning, or most of the work, is data curation." Husain, 10:42.
- "The more comprehensive your eval framework is, the cost of human review goes down, because you're automating more and more of these things and getting more confidence in your data." Husain, 11:12.
- "It's very very important to align the LLM judge to a human, because you need to know whether you can trust the LLM as a judge." Husain, 12:13.
- "I like to use a spreadsheet often, don't make it complicated." Husain, 12:45, on the alignment loop between the judge and a domain expert.
- "If you're having a conversation about evals and the first thing you start thinking about is tools, that's a smell that you're not going to be successful in your evaluations." Husain, 13:16.
- "You have to know what the process is before you jump straight into the tools, otherwise you're going to be blindsided." Husain, 13:47.
- "It's not that they're not valuable at all, it's just that you shouldn't rely on them because they can become a crutch." Husain, 14:18, on generic off the shelf evals.
- "If I'm looking at the data closely enough I can always find plenty of assertions and failure modes. It's not always the case, but it's often the case." Husain, 14:49.
- "Without the eval framework, a project similar to this seemed completely impossible for us." Sedgh, 15:19.
- "I wish we could just be a ChatGPT [wrapper] and manage to extract the experience we want for our users, but we never had that opportunity, because we had some really difficult cases." Sedgh, 15:52.
- "We never managed to get this working without fine tuning reliably." Sedgh, 16:23, on interleaving rendered interface elements with natural language in one response.
- "Creating something like this for a non savvy real estate agent may take a couple of hours to do, but using the agent they can do it in a minute." Sedgh, 17:57.
- "That essentially was not going to be possible without us using a comprehensive eval framework." Sedgh, 17:57.
- "Nailed the timing, thank you guys." 17:57, the line that closes the recording as Sedgh finishes with seconds to spare. The captions do not identify who says it.
Where this sits in the LLM Learning track
A Hackers' Guide to Language Models ends on the mistake almost everyone makes about their own evaluation: believing the system works because it worked on the examples you happened to choose. This page is the answer to that, and it is the least glamorous and most load bearing stretch of the track. Assertions in CI, human review in a tool you built so that reviewing is frictionless, then an LLM judge that has been measured against a human, and only then the fine tuning that the curated data makes possible.
Read it before Building Effective Agents with LangGraph. An agent multiplies the number of places a system can go wrong, and every one of those places is invisible without a trace viewer and a scoring loop. The two pages together are the shape of a real LLM product: a workflow you can draw, an eval system that ratchets, and an agent only where the path genuinely cannot be written down.
Resources mentioned
The talk and the speakers
- How to construct domain-specific LLM evaluation systems, the talk page, at AI Engineer World's Fair 2024 (San Francisco, 25 to 27 June 2024), part of the AI Engineer conference series
- Hamel Husain, founder of Parlance Labs (speaker page, organization page, @HamelHusain)
- Emil Sedgh, CTO at Rechat and an architect of Lucy (@emilsedgh)
- Rechat introduces Lucy, the September 2023 announcement of the assistant this talk is about
The written version of this material
- Your AI Product Needs Evals, Husain, 29 March 2024. The companion essay, with the Rechat case study in detail and the counts the talk omits. Every figure in the scaffolding section above comes from here
- Using LLM-as-a-Judge For Evaluation: A Complete Guide, Husain, October 2024. The depth on the topic the talk declares out of scope. (This post was previously titled "Creating a LLM-as-a-Judge That Drives Business Results," so older citations point at the same URL under the old name)
- A Field Guide to Rapidly Improving AI Products, Husain's later and broader treatment of the same loop
- AI Evals For Engineers & PMs, the course Husain teaches with Shreya Shankar
Tools named on stage
- LangSmith, the tracing tool Rechat chose, and its observability concepts
- Metabase, already in use at Rechat and therefore where the assertion results went; dashboards guide
- Shiny for Python (get started), Husain's choice for the custom annotation app, with Gradio and Streamlit named as equally good options
Background
- ReAct: Synergizing Reasoning and Acting in Language Models, the pattern the first prototype was built on
- GPT 3.5, the model behind that prototype


