youtube.nixfred.com nixfred.com
Creator

AI Engineer

Conference talks from practitioners shipping LLM products, heavy on the unglamorous parts that decide whether one works.

1video

← All videos

18:44
AI Engineer

How to Construct Domain Specific LLM Evaluation Systems: Hamel Husain and Emil Sedgh

Emil Sedgh, CTO at Rechat, and Hamel Husain of Parlance Labs walk through how Lucy, the assistant inside Rechat's real estate CRM, got from a GPT 3.5 and ReAct prototype that was slow, wrong most of the time and majestic when it worked, to a product whose success rate they could actually measure. The recipe in order: assertions written from failures observed in real traces and run in CI, results logged to the Metabase they already owned, traces in LangSmith, and a custom Shiny for Python trace viewer because the off the shelf tools carried too much friction. Synthetic inputs from a model roleplaying a real estate agent supply the test cases, plain prompt engineering is the smoke test for the pipeline itself, and only then an LLM judge, trusted exactly as far as its measured agreement with a domain expert. Husain names the four ways teams get this wrong; Sedgh closes with the three behaviors that needed fine tuning rather than clever prompting.

AIDevOpsSep 19, 2024