At a glance
In December 2024 Anthropic published Building Effective Agents, an engineering post that cut through a year of agent hype with one distinction: a workflow is scaffolding of predefined code paths around model calls, and an agent is what you get when you remove the scaffolding and let the model direct its own actions. This video is Lance Martin of LangChain opening a notebook and building every single pattern in that post from scratch in LangGraph, with the code on screen and a real run at the end of each one.
He builds eight things in thirty one minutes: the augmented model call with structured output, the same call with tools, then prompt chaining with a gate, parallelization, routing, orchestrator and workers, evaluator and optimizer, and finally a tool calling agent with no scaffolding left. Every one of them is the same object underneath, a StateGraph with a typed state container, nodes that are plain Python functions, and edges that are either fixed or conditional on what is in the state.
What makes it worth the half hour is that it is a design argument disguised as a code walkthrough. The patterns are ordered by how much control passes from your code to the model, and Lance is unusually honest about where the line sits today. He has shipped far more workflows than agents over the last two years, he says, because agents have not been reliable with large tool sets or long tool call trajectories, and the field currently prefers workflows in production precisely because the scaffolding is what makes them trustworthy. He also tells you up front that none of this needs a framework, and then makes the narrow case for one anyway.
What a workflow is, and what an agent is
Lance opens with the definition, and he draws it as a spectrum rather than a binary. His exact framing: think about a workflow as "some kind of scaffolding of predefined code paths around LLM calls." This has been popular for years, and in many cases it makes a lot of sense to take model calls and embed them in a fixed set of code paths.
Then he points at the middle of his own drawing. Sometimes you can have the model decide which paths to take inside a workflow, and that is a third category sitting between the two poles. The scaffolding still exists, you still wrote every path that can be taken, but the choice of which one runs is delegated to a model call. Routing, orchestrator and workers, and evaluator and optimizer all live in this middle band, and he flags it at 00:31 before he has shown a line of code, because it is the part people collapse when they argue about whether something is "really" an agent.
Agents remove the scaffolding. You let the model direct its own actions, and in this formulation actions are tool calls. The model receives the feedback from those tool calls and decides what to do next. He immediately heads off the obvious objection: you can certainly have models performing tool calls inside workflows, and plenty of workflows do. The difference is not the presence of tool calls. It is that a workflow has scaffolding around them and an agent is unbounded. An agent "just directly receives the output of environmental feedback from tool calls, decides what to do next, without the kind of reasoning scaffolding around it." That is the differentiation, and it is the spine of everything that follows.
Why a framework at all
Before any code, at 01:32, Lance stops to answer the objection he knows is coming. "It's important to call out why frameworks," he says. "Implementing these patterns doesn't require framework, right. Sometimes this can be done in a few lines of code. I completely understand that." That concession is not throat clearing. It is the premise of his actual claim, which is narrow.
LangGraph "really aims to minimize the overhead of implementing these patterns." And then the two things it explicitly does not do, which he repeats twice in thirty seconds because it is the whole pitch: it does not abstract prompts, and it does not abstract architecture. What is happening is that LangGraph is "supporting infrastructure underneath any workflow or agent." You write the prompts. You draw the graph. The framework sits below both.
The infrastructure is three things, and he names them in order:
- Persistence. This gives you memory. It also gives you human in the loop, "the ability to pause for example while an agent is processing and approve a tool call, pause and review what the agent is doing." He calls this extremely useful in many cases, and it is the item he comes back to in the closing minute.
- Streaming. Model APIs stream tokens, he grants, everyone knows that. But when you build these workflows and agents you often want to stream something else: what a particular step output, the output of a tool call. "You often want more flexibility over streaming than just what the LLM calls produce," and LangGraph gives you control over what you emit from any point in a run. He flags this as very useful specifically when putting these things in production.
- Deployment. Testing, debugging and deploying. "It is extremely easy to go from any workflow or agent you implement to a deployment."
Then the summary line, stated flatly at 03:06: "LangGraph doesn't abstract prompts and it does not abstract architecture. It's really giving you low level infrastructure that sits underneath any of these workflows or agents." Note what is absent from that pitch. There is no claim that the framework makes the patterns better, shorter or smarter. The claim is that the patterns are yours and the plumbing is theirs.
The building block: the augmented LLM
The post's first figure is the augmented model call, and Lance treats it as the atom every later pattern is assembled from. Models can be augmented with many things. Memory is one, and that is one of the things a framework like LangGraph provides. Models can also interact with tools, "and this is the foundation for building many workflows and agents."
So he builds it from scratch. The notebook setup is three lines of housekeeping: pip install a few packages, set the API key, create the model. The model he runs on through the whole video is Claude, which is visible at 08:42 when he narrates the chain's last step as "Claude creates final joke."
Augmentation one, structured output. Take a schema, in this case a Pydantic model, and bind it to the model. He uses LangChain's with_structured_output method to do it, and then immediately adds the caveat that matters for anyone evaluating the stack: "again, LangGraph does not require you to use LangChain. You can use the raw model APIs, it's completely fine." He runs the cell and the output adheres to the schema he passed.
# a schema, then a model that is contractually obliged to fill it
class SearchQuery(BaseModel):
search_query: str = Field(None, description="Query that is optimized for web search.")
justification: str = Field(None, description="Why this query is relevant to the user's request.")
structured_llm = llm.with_structured_output(SearchQuery)
structured_llm.invoke("How does Calcium CT score relate to high cholesterol?")
Augmentation two, tool calling. He defines a function multiply as a tool the model has access to, uses bind_tools to attach it, and invokes the result with an input likely to elicit a call. The tool call comes back: it takes the input and creates the arguments necessary to actually run the function.
def multiply(a: int, b: int) -> int:
return a * b
llm_with_tools = llm.bind_tools([multiply])
msg = llm_with_tools.invoke("What is 2 times 3?")
msg.tool_calls
Then the sentence that the last third of the video pays off, said casually at 04:38: "remember, when LLMs are creating tool calls they're really just giving you payloads to actually run that tool. You could then run the tool, pass the output of the tool back to the LLM, and really doing that loop gives you an agent." The agent is already fully specified, nine minutes before he builds it. Everything between here and there is the scaffolding you put around this block instead.
Pattern 1: prompt chaining
The intuition. Each model call processes the output of the previous one. That is the whole pattern.
When to reach for it. "When you've attested you can decompose into a few different LLM calls," in Lance's words, meaning when you have confirmed the task actually splits into a handful of steps rather than hoping it does. The post's figure is an input, call one, a gate on the output of call one, then call two, then call three.
The example he builds. A chain that takes a topic from the user, has the model make a joke, checks that the joke has a punchline, and then improves it twice with two subsequent model calls. Three calls, one gate.
Step one: the state container
This is the part he describes as the only real conceptual work. "All I need to do is basically define a container for everything that I want to modify over my workflow." A TypedDict with four keys: topic, which comes from the user; joke, the output of the first call; improved_joke, the output of the second; final_joke, the output of the third.
class State(TypedDict):
topic: str
joke: str
improved_joke: str
final_joke: str
The state is not bookkeeping. It is the contract between the nodes, and every pattern in the video begins by writing one.
Step two: draw it before you write it
A small aside with more weight than it looks: "when laying out workflows I actually like to draw them out anyway, so this is kind of cool that the blog post has all these laid out." The blog post's figures are not illustrations to him, they are the design step he would have done on paper. And the instruction that follows is the one to keep: "when you draw these out, think about for example each of these calls or steps is just a different function."
So for this workflow there are three functions, generate_joke, improve_joke and polish_joke, and they are all just model calls.
Step three: nodes read state and return state
"What's interesting is with LangGraph, this container or state that I define is passed into every one of these steps, and I can extract whatever I want from it." He reads state["topic"] to get the topic the user wrote into state, and writes back by returning a dict with the key he wants to populate. generate_joke returns {"joke": ...}. improve_joke populates improved_joke. polish_joke populates final_joke.
def generate_joke(state: State):
msg = llm.invoke(f"Write a short joke about {state['topic']}")
return {"joke": msg.content}
def improve_joke(state: State):
msg = llm.invoke(f"Make this joke funnier by adding wordplay: {state['joke']}")
return {"improved_joke": msg.content}
def polish_joke(state: State):
msg = llm.invoke(f"Add a surprising twist to this joke: {state['improved_joke']}")
return {"final_joke": msg.content}
"So what's happening is at each of these steps I'm making LLM calls and I'm populating my state, or this container in the workflow, with the output of each LLM call. That's really all that's going on." A node is a function from state to a partial state update. That sentence describes every node in the video.
Step four: the gate
The gate is not a node. It is a plain function that takes state and returns a string, and the string is what the graph routes on. His check: does the joke contain a question mark or an exclamation point. He is explicit that this is "just some arbitrary criteria," picked because it is cheap to demonstrate, and the function can return anything you want. He returns the two strings Pass and Fail.
def check_punchline(state: State):
if "?" in state["joke"] or "!" in state["joke"]:
return "Pass"
return "Fail"
"Now this can be used as conditional edge or gate in LangGraph." One wrinkle worth noticing, because it is visible on screen and he does not dwell on it: in his wiring, Pass continues into the improvement chain and Fail ends the run. So a joke that already looks like a punchline gets polished twice and a joke that does not gets dropped on the floor, which is the opposite of what a quality gate usually does. With an arbitrary criterion it does not matter for the demo, and the shape of the wiring is the point, but if you lift this code the branch labels are the first thing to think about.
Step five: compile the graph
With a state, three node functions and a gate function in hand, the graph itself is six lines of plumbing.
workflow = StateGraph(State)
workflow.add_node("generate_joke", generate_joke)
workflow.add_node("improve_joke", improve_joke)
workflow.add_node("polish_joke", polish_joke)
workflow.add_edge(START, "generate_joke")
workflow.add_conditional_edges(
"generate_joke", check_punchline, {"Pass": "improve_joke", "Fail": END}
)
workflow.add_edge("improve_joke", "polish_joke")
workflow.add_edge("polish_joke", END)
chain = workflow.compile()
He narrates it exactly in that order: take in the state, initialize the workflow, add the three steps, "and this is where I define the connectivity of my graph or workflow." Start goes to generate_joke. Then the conditional edge hangs off generate_joke and dispatches on the gate's return value. Then two fixed edges carry the rest. Compile, and "there we go, that's the exact same workflow that they drew out in the blog post, applied to joke generation."
Step six: run it
"Once I have this workflow I can very simply just run chain.invoke. That's a very simple way to invoke any chain, workflow or whatever you have in LangGraph." He passes a topic from the user, and the run does what the graph says: the model creates an initial joke, improves it, and Claude creates the final joke.
Pattern 2: parallelization
When to reach for it. Two cases, and he gives a real one from his own work for each. First, when you want multiple perspectives on a single task: "I've used this quite a bit with things like multi query RAG. If I have a question, I want to fan it out into like three or four different sub questions." Second, when independent tasks can be performed with different prompts or different models. "Lots of cases where you want to just parallelize tasks in workflows."
The example. Take a topic and create a joke, a story and a poem, all at once, then combine them.
The state. Same move as before, "just a container for everything I'm going to modify in my workflow": a topic from the user, then joke, story, poem, and combined_output to hold the aggregate.
class State(TypedDict):
topic: str
joke: str
story: str
poem: str
combined_output: str
The nodes. Three model calls and one aggregation. call_llm_1 writes a joke about the topic, call_llm_2 a story, call_llm_3 a poem, and aggregator reads all three out of state and joins them into a single string. "Super simple," he says, and it is: four functions, none longer than three lines.
The graph. This is the shortest build in the video, because the parallelism is not a feature you switch on. It is a consequence of the edges you drew.
parallel_builder = StateGraph(State)
parallel_builder.add_node("call_llm_1", call_llm_1)
parallel_builder.add_node("call_llm_2", call_llm_2)
parallel_builder.add_node("call_llm_3", call_llm_3)
parallel_builder.add_node("aggregator", aggregator)
parallel_builder.add_edge(START, "call_llm_1")
parallel_builder.add_edge(START, "call_llm_2")
parallel_builder.add_edge(START, "call_llm_3")
parallel_builder.add_edge("call_llm_1", "aggregator")
parallel_builder.add_edge("call_llm_2", "aggregator")
parallel_builder.add_edge("call_llm_3", "aggregator")
parallel_builder.add_edge("aggregator", END)
parallel_workflow = parallel_builder.compile()
"In this case I go from start to LLM call one, two, three, and they all connect to the aggregator on the end." Three edges out of START means three nodes with no dependency between them, so they run together. Three edges into aggregator means it waits for all three. He displays the graph image, calls it "a nice visualization of your workflow," runs it, and gets a story, a joke and a poem created in parallel.
What it buys you. Wall clock time in the sectioning case. In the voting case, which the post also covers, running the same task several times and aggregating buys reliability on judgment calls where a single sample is noisy.
Pattern 3: routing
When to reach for it. "I've used routing a lot and it's extremely useful in a bunch of use cases." The trigger is wanting to go to exactly one of several steps, and wanting the model to make that choice. His concrete example from production work: "I've used routing a lot in working with retrieval, in taking a question and routing it to different retrieval systems."
The example. Take an input and route it to joke, story or poem generation based on what the user actually asked for.
The trick he wants you to steal. "Now I'm going to show a very nice trick here for routing, I often like to do this." Give the routing model a structured output, so it is guaranteed to produce one of the three strings as a structured object rather than prose you have to parse. The schema is a Pydantic model with a single field typed as a Literal of the three options.
class Route(BaseModel):
step: Literal["poem", "story", "joke"] = Field(None, description="The next step in the routing process")
router = llm.with_structured_output(Route)
"That guarantees the LLM is going to produce, in this particular case, the strings poem, story or joke as a structured object." The guarantee is the point. Without it the router node becomes string matching against a sentence, and that is where routers break.
The state. An input from the user, a decision from the router, and an output saved as the final result.
The nodes. Three model calls that write a story, a joke or a poem, plus the router node. The router takes the user input, decides which of the three the content calls for, returns the Pydantic object, and then Lance does the one extra step people forget: he extracts .step off the returned model and writes it into state as decision.
def llm_call_router(state: State):
decision = router.invoke([
SystemMessage(content="Route the input to story, joke, or poem based on the user's request."),
HumanMessage(content=state["input"]),
])
return {"decision": decision.step}
def route_decision(state: State):
if state["decision"] == "story":
return "llm_call_1"
elif state["decision"] == "joke":
return "llm_call_2"
elif state["decision"] == "poem":
return "llm_call_3"
"So then I have my decision in state, and now this is just going to be a conditional edge that'll look at the decision and determine what node to go to. Super simple. So basically if the decision is story go to node one, joke two, poem three. That's it."
The honesty about the toy. He breaks frame deliberately here, because the example is degenerate: "you'll see in this toy example each of these steps are doing the same thing, but in real world examples the router will send to different steps that will have different logic, different LLM calls. This is just a toy example showing you how to hook up that logic between for example a router step, a structured output, and a decision about where to go next."
The rendering detail. When he visualizes the compiled graph he points at something concrete: "what's kind of nice in LangGraph, when we visualize this, this dotted line means a conditional edge. So it's going to go to only one of those three paths whenever this runs, and of course that's going to be based upon the decision of the router." A dotted edge in a LangGraph render is a promise about runtime, not a stylistic choice, and it is how you read someone else's graph at a glance.
The run. He adds a print statement to each of the three nodes so the output tells you which node was visited, runs it, confirms the joke node fired, and gets the joke back.
Pattern 4: orchestrator and workers
This is the one he spends the longest on, and the one with a genuinely new API in it.
What it is. "A case where you want an LLM to break down a task into a set of subtasks, delegate each subtask to an independent worker, and then synthesize the results."
How it differs from parallelization. This is the sentence to hold onto: "it's kind of like parallelization, except the key difference is this worker assignment you don't know ahead of time. So you're having an LLM reason about something and then create a bunch of workers based upon its reasoning." And then the structural consequence: "so again in this case the LLM is kind of gating or creating the control flow, just like in the case of routing." Parallelization has a fan-out width you typed into your graph. Here the width is an output of a model call.
The example. Report writing, which he says he has used a lot. "Maybe you've played with deep research. An LLM reasons about the plan for the report and dynamically generates a bunch of report sections, and then goes and does research on all of them. Classic example of an orchestrator worker type workflow." LangChain later shipped Open Deep Research as an open source build of exactly this shape, if you want to read a production sized version of the pattern.
The planner is structured output again
"For this, the trick is I'm also going to use structured outputs." Two Pydantic models: a Section with a name and a description, and a Sections wrapper holding a list of them. Bind the list schema to the model and that bound object is the planner.
class Section(BaseModel):
name: str = Field(description="Name for this section of the report.")
description: str = Field(description="Brief overview of the main topics and concepts to be covered in this section.")
class Sections(BaseModel):
sections: List[Section] = Field(description="Sections of the report.")
planner = llm.with_structured_output(Sections)
"So what's cool here is the planner is going to take an input, reflect on it, and produce a list of sections based upon its reflection. So this is dynamic, I don't know how many sections it'll create a priori. That's why this is a very good orchestrator worker use case." The test for whether you need this pattern is exactly that: if you can count the branches before the run, you do not need it.
Two states, not one
This is the part he flags as the real design wrinkle, and it is the only place in the video where a pattern needs more than one state object.
The orchestrator's graph state has four keys: topic from the user, sections (the planner's list), completed_sections (which all the workers write into), and final_report.
The workers get their own state. "This is where things get interesting with orchestrator worker workflows in LangGraph. The way we often like to do it is, for the workers, give them their own state." His reason: "each of those workers, you want to handle independent inputs, and they're kind of all self-contained objects. Think about them as their own little buckets in which work's being done, and different work's being done in each one, but they're all writing out to the same output."
class State(TypedDict):
topic: str
sections: list[Section]
completed_sections: Annotated[list, operator.add]
final_report: str
class WorkerState(TypedDict):
section: Section
completed_sections: Annotated[list, operator.add]
"And this is why I include this completed_sections key in the worker state and in the graph state." Then the mechanism: "what's interesting in LangGraph is when you have overlapping keys, and you write to for example this completed_sections key in each worker, the outer state will also have that update reflected."
The piece that makes the parallel writes safe is the annotation. "All the workers are going to write to this completed_sections key in parallel, and we structure this key with an annotation that allows for the addition of new elements." That is a reducer: Annotated[list, operator.add] tells LangGraph that two concurrent writes to this key should be concatenated rather than one overwriting the other. Without it, four workers finishing at once would leave you with one section. "So that's really all we need to do."
The three node functions
The orchestrator is the planner, invoked with the topic and told to create a plan. The worker takes WorkerState, is handed a section name and description, writes that one section, and returns it under completed_sections. The synthesizer reads completed_sections out of state, joins them into a string, and writes final_report.
def orchestrator(state: State):
report_sections = planner.invoke([
SystemMessage(content="Generate a plan for the report."),
HumanMessage(content=f"Here is the report topic: {state['topic']}"),
])
return {"sections": report_sections.sections}
def llm_call(state: WorkerState):
section = llm.invoke([
SystemMessage(content="Write a report section following the provided name and description."),
HumanMessage(content=f"Here is the section name: {state['section'].name} and description: {state['section'].description}"),
])
return {"completed_sections": [section.content]}
def synthesizer(state: State):
completed_report_sections = "\n\n---\n\n".join(state["completed_sections"])
return {"final_report": completed_report_sections}
"This completed_sections is a state key that all the workers can write to in parallel, that's the key point, and all those sections are going to be accumulated in that completed_sections key. Then I'm going to have a synthesizer that's going to read out the completed sections and just write them all out as a string. That's it."
The one genuinely new API: Send
"Now this is the only thing that's new and a little bit special in orchestrator worker style workflows. Because this is so common, we have a special API called send in LangGraph that allows you to basically spawn these workers dynamically."
The mechanism is small. The planner already wrote the sections into state. You write a function that iterates them and returns a list of Send objects, one per section, each naming the node to run and the state to initialize it with. You hang that function off the orchestrator as a conditional edge.
def assign_workers(state: State):
return [Send("llm_call", {"section": s}) for s in state["sections"]]
"For each section in sections, send to llm_call, that's my worker, and basically initialize the state as section when I do that. So then that llm_call receives worker state, and it's receiving from state the section name and description, and then it goes and writes that section. It writes its output to completed_sections, which my orchestrator has access to, and then the synthesizer basically just grabs completed sections and combines them. That's it, you're done, that's all you need to do."
orchestrator_worker_builder = StateGraph(State)
orchestrator_worker_builder.add_node("orchestrator", orchestrator)
orchestrator_worker_builder.add_node("llm_call", llm_call)
orchestrator_worker_builder.add_node("synthesizer", synthesizer)
orchestrator_worker_builder.add_edge(START, "orchestrator")
orchestrator_worker_builder.add_conditional_edges("orchestrator", assign_workers, ["llm_call"])
orchestrator_worker_builder.add_edge("llm_call", "synthesizer")
orchestrator_worker_builder.add_edge("synthesizer", END)
orchestrator_worker = orchestrator_worker_builder.compile()
Reading the rendered graph
Another concrete note on the visualization, and this one is useful: "this dotted line just shows you that you use the send API to spawn a whole bunch of llm_call workers. You don't know how many, a heap ton, that's why it doesn't draw out each one specifically. It's going to be dynamically determined based upon the orchestrator plan. That's the key characteristic of these orchestrator worker workflows: you don't know a priori how many workers you need, the LLM will determine that on the fly."
So in a LangGraph render, a single dotted edge into one worker node is not a single worker. It is an unknown number of them, and the picture cannot tell you more than that.
The run
He asks for a report on LLM scaling laws. "What's kind of cool is this runs fairly quickly." The planner generates the report sections, he opens final_report, and the sections are concatenated in plan order: "you can see I get this rich introduction, and then I get, you know, fundamentals, scaling relationships, following the plan."
Then the caveat, which is the kind of thing that separates a demo from a claim: "now in reality when I do this I have much more detailed prompts than I showed here. This is really showing you the workflow rather than the particulars of how I build a high quality report writer. In fact I have separate videos on report writing you could check out, but this is just showing you how to set up an orchestrator worker style of workflow." The graph shape is the deliverable. The prompt quality is a different, longer problem.
Pattern 5: evaluator and optimizer
What it is. One model generates a response, another grades it and gives feedback, in a loop. He introduces it as the sibling of the previous pattern: "if we just talked about the orchestrator worker workflow, let's talk about something that's a bit related, the evaluator optimizer workflow. In both these cases you're going to have LLMs directing the control flow through predefined code paths."
Where he has used it. Grading RAG outputs. "I've used this quite a bit, for example grading responses from a RAG system for hallucinations or for factual accuracy. I've used these kind of like evaluator gates, and if for example there's hallucination in the output that isn't grounded by the documents, I send it back and have it regenerate a response." That is the canonical use: the quality bar is checkable even when it is not reliably hittable on the first try.
The structured outputs aside, which is the best thirty seconds in the video
Before building it he stops to generalize, because this is the fourth pattern in a row that leans on the same primitive. "Here I'm going to use structured outputs again. You see there's kind of a trend here. Structured outputs is like, kind of, all you need, in quotes."
Then the argument: "it's extremely convenient because you can really build all these workflows just off structured outputs. You don't have to use tool calling, for example. You can use routers to basically conditionally determine where to go next, you can have then nodes that for example just call tools depending on the result of the router itself. So really, just structured outputs is a nice way you can build a lot of complex workflows."
That is a real engineering position and worth stating plainly: for workflows, a model that reliably fills a schema is enough. Tool calling is a capability you need for agents, not a prerequisite for scaffolding. Four of the five workflow patterns in this video route on a schema field, not a tool call.
The grader schema
class Feedback(BaseModel):
grade: Literal["funny", "not funny"] = Field(description="Decide if the joke is funny or not.")
feedback: str = Field(description="If the joke is not funny, provide feedback on how to improve it.")
evaluator = llm.with_structured_output(Feedback)
"That's going to be my my grader model, that's going to be grade and feedback. Decide if the joke is funny or not, in my case, and if it is not funny give some feedback."
The state and the loop
The state carries a topic, the joke, the feedback string and the funny_or_not grade. He describes the cycle in one breath: "I'm going to take a joke based upon a topic from a user, I'll generate the joke, and then I'll grade it, give it feedback, determine if it's funny or not. Based on the feedback I'll go back, regenerate a new joke. That's it, nice and simple."
The generator has the one piece of logic that makes the loop work, and he flags it himself: "you'll see I do something kind of interesting here. I'm going to check if there's feedback in the state. Now there might be, because I've basically looped back to this node if I determine that the joke is not good. So there may be feedback in the state. If there's feedback, I include it in my prompt, otherwise I just say write a joke about the topic."
def llm_call_generator(state: State):
if state.get("feedback"):
msg = llm.invoke(f"Write a joke about {state['topic']} but take into account the feedback: {state['feedback']}")
else:
msg = llm.invoke(f"Write a joke about {state['topic']}")
return {"joke": msg.content}
def llm_call_evaluator(state: State):
grade = evaluator.invoke(f"Grade the joke {state['joke']}")
return {"funny_or_not": grade.grade, "feedback": grade.feedback}
def route_joke(state: State):
if state["funny_or_not"] == "funny":
return "Accepted"
elif state["funny_or_not"] == "not funny":
return "Rejected + Feedback"
One node, two prompts, branching on whether a state key is populated. That is the entire difference between the first pass and every pass after it, and it is why the loop does not just regenerate the same joke forever.
The conditional edge that closes the cycle
"Then I have a conditional edge that'll look at state funny or not, so again that's like kind of my grade, and if it's funny route to accepted, if it's not funny route to rejected and feedback."
optimizer_builder = StateGraph(State)
optimizer_builder.add_node("llm_call_generator", llm_call_generator)
optimizer_builder.add_node("llm_call_evaluator", llm_call_evaluator)
optimizer_builder.add_edge(START, "llm_call_generator")
optimizer_builder.add_edge("llm_call_generator", "llm_call_evaluator")
optimizer_builder.add_conditional_edges(
"llm_call_evaluator",
route_joke,
{"Accepted": END, "Rejected + Feedback": "llm_call_generator"},
)
optimizer_workflow = optimizer_builder.compile()
"If accepted we go to end, if it is rejected and feedback we go back to llm_call_generator. So you can see, when you set the conditional edge, this is how you can basically route from the output of your logic, your edge logic, to the next node to go to. That's it, it's extremely simple, and you can see you get a nice visualization of it."
This is the first graph in the video with a cycle in it, and the cycle is not a special construct. It is a conditional edge whose target happens to be a node that already ran. That is the thing a graph gives you that a chain does not.
The run
Input: cats. He gets the joke out, then inspects the state to see what the grader said. state["feedback"]: "in this case it seems to like it." state["funny_or_not"]: funny. "So the grader liked the joke, we went ahead and returned it to the user, we ended. So we can see that the grader was initiated and decided to like the joke, so it passes and we finish." One pass, no revision, which is a slightly anticlimactic demo of a loop and he does not pretend otherwise.
Removing the scaffolding
At 23:32 he draws the line under the workflow half of the video and restates what has actually been happening: "we talked about a number of different workflows that use LLMs within some kind of reasoning scaffolding, and in the case of orchestrator worker, evaluator optimizer, routing, you actually do let the LLM make decisions to route the control flow through that scaffolding."
That is the middle band of Figure 1, named explicitly, with its three members listed. Three of the five workflows already hand control flow decisions to a model. What makes them workflows is that every path the model can choose is a path he typed.
"Now let's remove the scaffolding and let's talk about agents. With agents you're simply allowing an LLM to perform actions in the form of tool calls, and directly receive the output or feedback from those actions. And so in the workflow case we talked about, there are always kind of these predefined code paths that we had an LLM follow and route through. In the case of an agent, we've removed those."
When do you actually need one
"This is kind of the big question." His answer has two halves, and the second half is the more useful one.
The case for. "You see agents being used in cases where you really have open ended problems that you cannot easily capture in a workflow. For example, you want the LLM to utilize different tools in a pattern that you just cannot predict a priori, so it's not easy to lay it out in the workflow." The existence proof he cites is SWE-bench, "a benchmark for software engineering in which Anthropic actually used an agent architecture, just shown like this, and achieves very strong performance." That is a real result: Anthropic's write-up on SWE-bench Verified describes exactly this shape, a model in a loop with bash and file editing tools and very little scaffolding. "So we know for certain open ended tasks, agents are appropriate."
The caveat he puts right next to it. "I do want to caveat, in the event that LLMs get extremely proficient at tool calling, it's also possible that a lot of the scaffolding that we talked about with various workflows is unnecessary. Today, if you know roughly the sequence tools need to be initiated, it's often better just capture it in a workflow, in terms of reliability, than just give it to an agent and let the agent hopefully call that correct sequence."
Read those two paragraphs together and you have the decision rule the whole video is built around. The question is not how smart the model is. It is whether you know the sequence. If you know it, write it down, because a written sequence runs the same way every time and a hoped-for sequence does not.
Pattern 6: the agent loop, built from scratch
"Now let's just set up an LLM with tools." Three tools, all arithmetic: multiply, add and divide. "Nice and simple."
def multiply(a: int, b: int) -> int:
return a * b
def add(a: int, b: int) -> int:
return a + b
def divide(a: int, b: int) -> float:
return a / b
tools = [add, multiply, divide]
tools_by_name = {tool.name: tool for tool in tools}
llm_with_tools = llm.bind_tools(tools)
The state is one key
The whole agent runs on a single state key. "My state has a single key, in this particular case messages, which is going to accumulate what I pass as the user, what the LLM produces, and so forth." That is MessagesState, a prebuilt state whose one key is a message list with an append reducer on it, the same mechanism as completed_sections in the orchestrator but applied to a conversation instead of report fragments.
Two nodes
llm_call. The model with tools bound, given a system message: "you're a helpful assistant tasked with performing arithmetic." Its output is appended to messages.
tool_node. "It's going to look at the state, look at the last message, determine if it's a tool call. If it is, it'll go ahead and just call that tool. That's it, nice and simple. And it's going to return that to state."
def llm_call(state: MessagesState):
return {"messages": [llm_with_tools.invoke(
[SystemMessage(content="You are a helpful assistant tasked with performing arithmetic on a set of inputs.")]
+ state["messages"]
)]}
def tool_node(state: dict):
result = []
for tool_call in state["messages"][-1].tool_calls:
tool = tools_by_name[tool_call["name"]]
observation = tool.invoke(tool_call["args"])
result.append(ToolMessage(content=observation, tool_call_id=tool_call["id"]))
return {"messages": result}
Environmental feedback, defined
This is the clearest definition of the term anywhere in the video, and he gives it while tracing the message list: "I'm going to have a sequence of human input, model, agent in this case decides to call a tool, this tool_node looks and sees, oh, the LLM decided to call a tool. It actually runs that tool call, that then is written to state messages as a tool message. This is that environmental feedback thing you hear about. So you talk about agents, they can perform actions, that's the tool call done up here, fine. They also can receive feedback from the environment and act on it. That is the output of this tool node. This tool node is basically the environmental feedback saying here's the output of the tool call. The LLM then will get that, decide what to do next. That's it, that's an agent."
The term stops being mystical once you see it is a ToolMessage appended to a list.
One conditional edge
"The only other thing I need is this conditional edge that basically says, was the last message a tool call? If so I'm going to go ahead and route to the tool node, and if not I will end. You can modify this in different ways, but a lot of times people basically just say allow the agent to continue making tool calls until it decides it doesn't need one anymore, and then you're done."
def should_continue(state: MessagesState) -> Literal["tool_node", END]:
last_message = state["messages"][-1]
if last_message.tool_calls:
return "tool_node"
return END
agent_builder = StateGraph(MessagesState)
agent_builder.add_node("llm_call", llm_call)
agent_builder.add_node("tool_node", tool_node)
agent_builder.add_edge(START, "llm_call")
agent_builder.add_conditional_edges("llm_call", should_continue, ["tool_node", END])
agent_builder.add_edge("tool_node", "llm_call")
agent = agent_builder.compile()
Count the pieces. One state key, two nodes, one fixed edge back from the tool node to the model, one conditional edge out of the model. That is the whole agent, and the loop exists because tool_node points back at llm_call.
"So there we are, I mean that's our agent loop, and that's kind of why agents are elegant. They're extremely simple in this kind of formulation. All it is, is basically an LLM initially with a bunch of tools, and it's like a tool node that will run the tool for you and return that environmental feedback to the LLM, and let the LLM keep spinning until it decides it doesn't need a tool call anymore, then you're done. That's really it."
Who owns the actions
A small but load bearing point about the division of labor: "now again, this tool call thing, think about that as actions, so this LLM can perform actions. The actions are determined by the user, so you can pass in any set of tools to this LLM, it decides to perform those actions. Now you need some system that actually does those actions, so that's what we create with this tool node. And this just loops until the LLM says I don't need a tool call anymore, and then that conditional edge says okay, I can just end."
You choose the action space. The model chooses the sequence. Your code executes. That split is the entire security and reliability surface of an agent, stated in one paragraph.
The run
"I'm going to basically tell this agent: add three and four, then take the output and multiply by four."
And then the self aware bit, because the demo is trivially solvable without tools: "look, this is extremely simple and an LLM can just do this, I totally get it. This is more showing you the principle: setting up an agent with these tools for addition and multiplication, and testing whether or not it can correctly perform these tool calls in sequence, and looking at the flow of messages."
The flow of messages is the deliverable, and he walks it: "you have an input from the human, my instructions. The LLM looks at that and says, okay, I need to make a tool call, makes a tool call. The tool node executes the tool call, returns environmental feedback to my LLM as the result of the tool call, which is seven. My agent thinks about it, makes another tool call, responds with twenty eight. Agent thinks about it, final result is twenty eight, here's how we got there, no more tool calls needed, done."
Then the line the video is remembered for, at 28:35: "for all the hype agents, that's all it is: extremely simple, is tool calling in a loop."
So why not use agents for everything
Having just called the agent elegant, he spends the next minute arguing against reaching for it, and this is the most quotable stretch of the video because it is a practitioner reporting what he actually ships.
"Now again, why don't we just do this? Because look, a lot of problems can be solved with workflows, which are a bit simpler. Agents haven't been particularly reliable to date, particularly with large numbers of tools or complex trajectories of tool calls, and so a lot of people actually in production prefer workflows. I've done a lot more workflows than agents, to be honest, over the last year or two."
Two specific failure conditions are named there, and they are worth separating because they fail differently. Large numbers of tools is a selection problem: the model has to pick correctly from a wide menu on every turn. Complex trajectories of tool calls is a compounding problem: a long sequence multiplies per step error rates until the run ends somewhere unrelated to the goal. A workflow removes the first by constraining the menu per node and removes the second by fixing the order.
Then the forecast, hedged precisely: "I completely acknowledge, in the event that you have very capable reasoning models that can perform low latency high quality tool calling, and we have more confidence that they can perform reliably in production, I think you will see the movement to this classic style of very simple tool calling agent in production. But the game in the field today is a little bit more like, people are saying, well, I want to put workflows into production because they're a little bit more trustworthy. Again, I have some scaffolding around the core LLM calls."
Note what is doing the work in that conditional. Not capability, but confidence. The blocker is not that models cannot call tools, it is that teams cannot yet predict how a given agent will behave on inputs they have not seen, and scaffolding is how you buy predictability back.
The prebuilt shortcut
At 29:35 he opens the documentation and shows the thing he deliberately did not use: "I do also want to show, this is looking at the documentation, that we do have a prebuilt method called create_react_agent that basically wraps what we just built from scratch right here, for convenience."
And he refuses to have an opinion about which you should use: "now this depends. If you want to build it from scratch yourself just as we did, that is completely fine. If you want to use the prebuilt method, that's fine as well, but it's available to you."
He also commits to sharing the notebook: "I will be sharing the link to this tutorial page which has everything we just went through, so all the different workflows we talked through are all in this tutorial." That page is Workflows and agents in the LangGraph docs, and every snippet above is in it, in Python and TypeScript, which is the single most useful artifact attached to this video.
The framework argument, restated with the code behind it
The last ninety seconds return to the three claims from minute two, now that you have watched eight graphs get compiled. "If I go back to the why LangGraph story, when you compile those workflows or agents in LangGraph, you're getting a persistence layer for free."
Persistence. "Which gives you short and long term memory, and that also gives you the ability to stop, perform interruptions, review and continue, IE human in the loop." Three capabilities from one mechanism: memory across turns, resumption after a stop, and a review gate in the middle of a run.
Streaming. "It gives you a whole bunch of streaming capacities. You can stream independent values from your state at any point in time, you can stream of course tokens out of LLM calls." The first half is the one that is hard to get otherwise. Streaming tokens is a model API feature; streaming an arbitrary state key as a node writes it is a framework feature, and it is what makes a multi minute orchestrator run feel alive rather than hung.
Deployment. "We have a very nice and easy onramp for testing, debugging and deploying any of these, so you could take any of those workflows we just built and deploy them in like five minutes. So very, very quickly, that's really what you're getting with the framework."
He closes on the direction of travel, which is also an admission that the overhead is real: "and you can also see with LangGraph it's pretty easy to lay them out, and we're actually working on making it even easier. So we're trying to reduce the overhead, so you can lay it out almost as if you're writing Python, you're not even thinking about the framework, but you're getting these benefits kind of for free when you do use a framework. That's kind of the big idea here."
Then: "hopefully this was a helpful overview to present how to lay these various workflows or agents out using LangGraph, and what benefits you may get from LangGraph as a consideration. So thanks very much, feel free to leave any comments below."
Every pattern, as a reference table
| Pattern | What it is | Who picks the path | LangGraph primitive | Reach for it when | Wrong choice when |
|---|---|---|---|---|---|
| Augmented LLM | One model call with structured output, tools, retrieval or memory attached | your code | with_structured_output, bind_tools | Always. It is the atom every pattern below is built from | Never, but on its own it has no control flow at all |
| Prompt chaining | A fixed sequence of calls, each consuming the last one's output, with an optional gate partway | your code | add_edge in a line, plus one add_conditional_edges for the gate | You have confirmed the task decomposes into a few steps, and each step gets a simpler instruction | The steps are independent. You are paying serial latency for nothing |
| Parallelization | Several independent calls at once, joined by an aggregator. Sectioning or voting | your code | Multiple add_edge calls out of START, all converging on one node | You want several perspectives on one task, or independent subtasks with different prompts or models | A later branch needs an earlier branch's output. The join waits for everything |
| Routing | Classify the input, then run exactly one specialized branch | a model call, from a fixed menu | add_conditional_edges on a Literal field from structured output | One prompt is trying to cover several different jobs and doing all of them adequately | You can classify with a regular expression. The model call is pure cost |
| Orchestrator and workers | A planner decides the subtasks, workers run them in parallel, a synthesizer combines | a model call, menu and width | Send from a conditional edge, plus Annotated[list, operator.add] on the shared key | You genuinely cannot count the branches before the run. Report writing, deep research | You can count them. Use parallelization and keep the width in your source code |
| Evaluator and optimizer | A generator and a critic in a cycle, revising until the grade passes | a model call, loop or exit | A conditional edge pointing back at a node that already ran | Quality is easier to check than to achieve first try. Hallucination grading, style constraints | There is no checkable criterion, or the check costs as much as the generation. Also uncapped by default |
| Agent | A model with tools in a loop, choosing each next action from environmental feedback | the model, unbounded | MessagesState, a tool node, and one conditional edge on last_message.tool_calls | The sequence of tools genuinely cannot be predicted. Open ended coding, SWE-bench style tasks | You roughly know the sequence. Write it down instead, for reliability. Also unreliable with many tools or long trajectories |
Key takeaways
- A workflow is scaffolding of predefined code paths around model calls. An agent removes the scaffolding and lets the model direct its own actions. Tool calls appear in both, so their presence proves nothing.
- There is a third band between the two, and Lance draws it deliberately: a workflow where a model call decides which of your predefined paths to take. Routing, orchestrator and workers, and evaluator and optimizer all live there.
- Every pattern is the same augmented model call, wired differently. Nothing after the first five minutes of the video adds capability, only control flow.
- The LangGraph pitch is narrow and worth holding them to: it does not abstract prompts and it does not abstract architecture. It supplies persistence, streaming and deployment underneath a graph you drew yourself.
- Every build starts the same way: define a state container for everything the workflow will modify. Nodes are functions from state to a partial state update. Gates are functions from state to a label and write nothing.
- Structured output is the workhorse. Four of the five workflow patterns route on a schema field rather than a tool call, and Lance says outright that you can build most of this out of structured outputs alone.
- Orchestrator and workers is the one pattern with a new API in it.
Sendturns a planner's list into a dynamic number of worker invocations, andAnnotated[list, operator.add]is what lets those workers write to one state key concurrently without clobbering each other. - A cycle is not a special construct. It is a conditional edge pointing at a node that already ran, which is the thing a graph gives you that a chain does not.
- In a rendered LangGraph diagram, a dotted edge means exactly one path is taken at runtime, and a single dotted edge into a worker node can mean any number of workers.
- The agent is two nodes and one conditional edge. Environmental feedback is a
ToolMessageappended to a list. "For all the hype agents, that's all it is: extremely simple, is tool calling in a loop." - The decision rule: if you roughly know the sequence of tools, capture it in a workflow, because reliability beats flexibility. Build an agent when the sequence genuinely cannot be written down.
- Lance's own report from production, as of January 2025: he has shipped far more workflows than agents over the previous two years, agents have not been reliable with large tool sets or long tool call trajectories, and the field prefers workflows because the scaffolding is what makes them trustworthy.
Chapters
- 0:00:00 Introduction & Key Concepts
- 0:01:00 Understanding Workflows vs Agents
- 0:02:00 Why Use Frameworks? Benefits of LangGraph
- 0:04:00 Building Block: Augmented LLM
- 0:05:00 Pattern 1: Basic Prompt Chaining
- 0:09:00 Pattern 2: Parallelization
- 0:11:00 Pattern 3: Routing with LLMs
- 0:14:00 Pattern 4: Orchestrator-Worker Pattern
- 0:20:00 Pattern 5: Evaluator-Optimizer Workflow
- 0:24:00 Building Agents: Beyond Workflows
- 0:27:00 Implementing a Basic Agent Loop
- 0:30:00 Conclusion & LangGraph Benefits
Notable quotes
Think about a workflow as some kind of scaffolding of predefined code paths around LLM calls. Lance Martin, the definition the whole video runs on, 0:20
Agents remove the scaffolding, so you're basically letting an LLM direct its own actions. Lance Martin, the other half of the definition, 1:02
Workflows have some scaffolding around them, whereas an agent is unbounded. Lance Martin, after granting that workflows call tools too, 1:25
Implementing these patterns doesn't require framework, right? Sometimes this can be done in a few lines of code. I completely understand that. Lance Martin, opening the case for a framework by conceding the case against, 1:40
LangGraph doesn't abstract prompts and it does not abstract architecture. It's really giving you low level infrastructure that sits underneath any of these workflows or agents. Lance Martin, the pitch, stated twice in thirty seconds, 3:06
Remember, when LLMs are creating tool calls, they're really just giving you payloads to actually run that tool. You could then run the tool, pass the output of the tool back to the LLM, and really doing that loop gives you an agent. Lance Martin, specifying the agent nineteen minutes before building it, 4:40
All I need to do is basically define a container for everything that I want to modify over my workflow. Lance Martin, on state, the first step of every build in the video, 5:40
When laying out workflows I actually like to draw them out anyway, so this is kind of cool that the blog post has all these laid out. Lance Martin, on why the post's figures are a design step and not decoration, 5:58
You don't know a priori how many workers you need. The LLM will determine that on the fly. Lance Martin, on the one test for whether you need orchestrator and workers, 18:53
Structured outputs is like, kind of, all you need, in quotes. Lance Martin, after using it in four patterns running, 19:55
This is that environmental feedback thing you hear about. Lance Martin, pointing at a ToolMessage in a list, 26:10
That's it. That's an agent. Lance Martin, having written two node functions, 26:35
For all the hype agents, that's all it is: extremely simple, is tool calling in a loop. Lance Martin, after the arithmetic run returns 28, 28:50
Agents haven't been particularly reliable to date, particularly with large numbers of tools or complex trajectories of tool calls, and so a lot of people actually in production prefer workflows. Lance Martin, the honest state of the field in January 2025, 29:05
I've done a lot more workflows than agents, to be honest, over the last year or two. Lance Martin, reporting on his own work, 29:15
If you know roughly the sequence tools need to be initiated, it's often better just capture it in a workflow, in terms of reliability, than just give it to an agent and let the agent hopefully call that correct sequence. Lance Martin, the decision rule, 24:45
You could take any of those workflows we just built and deploy them in like five minutes. Lance Martin, closing on the third framework claim, 31:08
What the captions mangled
The auto captions on this video are rough on proper nouns, so a single correction pass rather than quiet edits throughout. Everything below is corrected in the quotes and prose above.
- LangChain appears as "Lang chain" and "landcraft". LangGraph appears as "Lang graph", "L graph", "langra", "lra" and "Lang craft".
- Anthropic appears once as "thropic", in the opening sentence.
- Pydantic appears as "pantic model" and "pedantic model" throughout. The library is Pydantic, and the method is
with_structured_output, rendered in the captions as "wi structured output". - SWE-bench appears as "s bench". It is SWE-bench, the software engineering benchmark, introduced in 2023.
- Conditional edge appears as "conditional ledge", "initial ledge" and "additional Edge" in several places. There is only one concept there.
- LLM scaling laws, his demo report topic, is transcribed as "L scaling logs".
- Smaller ones: "augment to LM" and "augmented elementer" are the augmented LLM, "prom chaining" is prompt chaining, "a HEB ton" is "a heap ton", "ver capacity reasoning models" is "very capable reasoning models", and "greater" is "grader" in the evaluator section.
Resources mentioned
The source material
- Building Effective Agents, the Anthropic engineering post this video implements pattern for pattern. Every figure Lance builds from is in it, and the workflow versus agent definition is theirs, not LangGraph's.
- Workflows and agents, the LangGraph tutorial page Lance promises to share at 30:07. It holds all eight builds in Python and TypeScript. Note that it has moved since the video: the URL on screen was under
langchain-ai.github.io/langgraph, which now redirects to thedocs.langchain.comaddress above.
The framework and the primitives he uses
- LangGraph and its source on GitHub
- The Graph API:
StateGraph,add_node,add_edge,add_conditional_edges,compile, and the reducer annotation that makes concurrent writes to one state key safe Send, the API for spawning a dynamic number of workers from a conditional edgecreate_react_agent, the prebuilt that wraps the agent he builds by hand- Persistence and checkpointers, memory, interrupts and human in the loop, and durable execution
- Streaming, including streaming arbitrary state values and not just model tokens
- Deployment and LangGraph Studio for the testing and debugging onramp he mentions
- LangChain itself, specifically structured output, tools and model bindings. He is explicit that LangGraph does not require any of it and raw model APIs are fine.
- LangSmith, the tracing and evaluation side of the stack, which is where you would actually look at one of these runs step by step
- Pydantic, the schema library behind every structured output in the video, plus Python's
TypedDictfor the state containers andoperator.addas the list reducer
Things he references in passing
- Claude, the model running every call in the notebook
- SWE-bench, the software engineering benchmark he cites as proof that agents are right for some open ended tasks, and Anthropic's SWE-bench Verified write-up describing the minimal agent architecture he is pointing at
- Multi query retrieval, his own stated use of parallelization: one question fanned into three or four sub questions
- Open Deep Research, LangChain's open source build of the orchestrator and workers report writer he demos, if you want the production sized version
- LLM scaling laws, the topic he types into the report writer demo
- ReAct: Synergizing Reasoning and Acting in Language Models, the paper the
create_react_agentname comes from - LangGraph Academy, the free course covering the same primitives at more length
- Lance Martin, the presenter, on X, and the LangChain YouTube channel where the separate report writing videos he mentions at 19:23 live
An honest footnote
Two things to hold alongside this video, neither of which undercuts it.
The forecast is the part to watch. Lance's own hedge at 24:33 and again at 29:05 is that the workflow preference is a statement about today, not about the shape of the problem: if models become reliable enough at low latency tool calling, much of the scaffolding becomes unnecessary and the simple tool calling agent moves into production. He recorded that in January 2025. That is a falsifiable prediction about a moving target, and it is the lens to read the video through rather than a timeless law. The patterns themselves do not expire, because the question they answer, do you know the sequence or not, does not depend on model capability. The default answer does.
The vendor position is stated, not hidden. This is a LangChain engineer making a case for LangGraph, and the case is unusually narrow: persistence, streaming, deployment, with prompts and architecture left to you. He opens by conceding that the patterns need no framework at all. Hold him to the narrow version. If you do not need checkpointing, human interrupts, cycles or streamed intermediate state, the five workflow patterns really are a few functions and some control flow, and the honest reading of his own argument is that you should write them that way.
A practical note on the code itself, since this page is meant to be usable as a reference: none of the eight builds has a step cap, a cost ceiling or a retry limit. The evaluator loop exits only when the grader says funny, and the agent loop exits only when the model stops asking for tools. That is correct for a teaching notebook and not correct for anything you point at production, and the bounding is on you because, as he says, the framework does not abstract architecture.
Where this sits in the LLM Learning track
This is the last of the four videos in Part 4 of the LLM Learning track, the stage about shipping something, and it closes the building half of the curriculum.
It reads directly out of the two videos before it. A Hackers' Guide to Language Models walks the practical ladder up to function calling and retrieval, which is the capability this video assumes on page one: every pattern here is built on a model that can fill a schema and emit a tool payload. How to Construct Domain Specific LLM Evaluation Systems is the real prerequisite, and the order matters more here than anywhere else in the track. Every pattern on this page multiplies the number of ways a system can fail, and a cycle or a dynamic fan-out multiplies it without bound. Lance's own reason for preferring workflows, that agents have not been reliable with long tool call trajectories, is a claim you can only make about your own system if you are already looking at traces and scoring runs. Read the evals page first, then read this one, and the pairing gives you the shape of a real product: a graph you can draw, an evaluation loop that ratchets, and an agent only where the sequence genuinely cannot be written down.
It also sets up Part 5. Andrej Karpathy's Software Is Changing (Again) argues for partial autonomy products with a human in the loop and an autonomy slider, and insists on the decade of agents rather than the year of them. That is the same argument from the product side that Lance makes from the implementation side, and the two videos agree on the mechanism: the slider Karpathy draws is the spectrum in Figure 1, and the thing that lets you move it is the interrupt and checkpoint machinery Lance names in his last ninety seconds. Ilya Sutskever on a decade of sequence to sequence learning is the research side of the same question, since whether Lance's forecast lands depends on whether the capability curve he is betting on keeps bending.
If you want one line to carry out of Part 4: reach for a workflow first, and build an agent only when the path genuinely cannot be written down.


