18:44How to Construct Domain Specific LLM Evaluation Systems: Hamel Husain and Emil Sedgh
Emil Sedgh, CTO at Rechat, and Hamel Husain of Parlance Labs walk through how Lucy, the assistant inside Rechat's real estate CRM, got from a GPT 3.5 and ReAct prototype that was slow, wrong most of the time and majestic when it worked, to a product whose success rate they could actually measure. The recipe in order: assertions written from failures observed in real traces and run in CI, results logged to the Metabase they already owned, traces in LangSmith, and a custom Shiny for Python trace viewer because the off the shelf tools carried too much friction. Synthetic inputs from a model roleplaying a real estate agent supply the test cases, plain prompt engineering is the smoke test for the pipeline itself, and only then an LLM judge, trusted exactly as far as its measured agreement with a domain expert. Husain names the four ways teams get this wrong; Sedgh closes with the three behaviors that needed fine tuning rather than clever prompting.