youtube.nixfred.com nixfred.com

Building Effective Agents with LangGraph

LangChain takes the patterns from Anthropic's Building Effective Agents post and implements every one of them from scratch in LangGraph: the augmented LLM, prompt chaining, parallelization, routing, orchestrator and workers, evaluator and optimizer, and finally a real agent looping over tools. The useful part is the boundary it draws. Workflows are predefined code paths and you should reach for them first; an agent is what you build when the sequence of steps genuinely cannot be known ahead of time.

Published Jan 27, 2025 31:49 video 60 min read Added Jul 30, 2026 Open on YouTube →

At a glance

In December 2024 Anthropic published Building Effective Agents, an engineering post that cut through a year of agent hype with one distinction: a workflow is scaffolding of predefined code paths around model calls, and an agent is what you get when you remove the scaffolding and let the model direct its own actions. This video is Lance Martin of LangChain opening a notebook and building every single pattern in that post from scratch in LangGraph, with the code on screen and a real run at the end of each one.

He builds eight things in thirty one minutes: the augmented model call with structured output, the same call with tools, then prompt chaining with a gate, parallelization, routing, orchestrator and workers, evaluator and optimizer, and finally a tool calling agent with no scaffolding left. Every one of them is the same object underneath, a StateGraph with a typed state container, nodes that are plain Python functions, and edges that are either fixed or conditional on what is in the state.

What makes it worth the half hour is that it is a design argument disguised as a code walkthrough. The patterns are ordered by how much control passes from your code to the model, and Lance is unusually honest about where the line sits today. He has shipped far more workflows than agents over the last two years, he says, because agents have not been reliable with large tool sets or long tool call trajectories, and the field currently prefers workflows in production precisely because the scaffolding is what makes them trustworthy. He also tells you up front that none of this needs a framework, and then makes the narrow case for one anyway.

What a workflow is, and what an agent is

Lance opens with the definition, and he draws it as a spectrum rather than a binary. His exact framing: think about a workflow as "some kind of scaffolding of predefined code paths around LLM calls." This has been popular for years, and in many cases it makes a lot of sense to take model calls and embed them in a fixed set of code paths.

Then he points at the middle of his own drawing. Sometimes you can have the model decide which paths to take inside a workflow, and that is a third category sitting between the two poles. The scaffolding still exists, you still wrote every path that can be taken, but the choice of which one runs is delegated to a model call. Routing, orchestrator and workers, and evaluator and optimizer all live in this middle band, and he flags it at 00:31 before he has shown a line of code, because it is the part people collapse when they argue about whether something is "really" an agent.

Agents remove the scaffolding. You let the model direct its own actions, and in this formulation actions are tool calls. The model receives the feedback from those tool calls and decides what to do next. He immediately heads off the obvious objection: you can certainly have models performing tool calls inside workflows, and plenty of workflows do. The difference is not the presence of tool calls. It is that a workflow has scaffolding around them and an agent is unbounded. An agent "just directly receives the output of environmental feedback from tool calls, decides what to do next, without the kind of reasoning scaffolding around it." That is the differentiation, and it is the spine of everything that follows.

YOUR CODE OWNS THE CONTROL FLOW THE MODEL OWNS THE CONTROL FLOW WORKFLOW, FIXED PATH Scaffolding you wrote Every step and its order is in your code. The model only fills in content at each step. chaining, parallelization WORKFLOW, MODEL ROUTES Scaffolding, model picks LLM You still wrote every path. A model call decides which one runs, or how many run. routing, orchestrator, evaluator AGENT No scaffolding at all LLM tools Unbounded. The model receives environmental feedback and chooses the next action itself. tool calling in a loop Tool calls appear in all three bands. What changes left to right is who decides the sequence. Lance draws the middle band explicitly, because it is the one people collapse into the other two.
Figure 1. The spectrum Lance sketches in the first ninety seconds. The presence of tool calls does not make something an agent; the absence of a path you wrote does. The five workflow patterns in this video occupy the left and middle bands, and only the last build in the video sits on the right.

Why a framework at all

Before any code, at 01:32, Lance stops to answer the objection he knows is coming. "It's important to call out why frameworks," he says. "Implementing these patterns doesn't require framework, right. Sometimes this can be done in a few lines of code. I completely understand that." That concession is not throat clearing. It is the premise of his actual claim, which is narrow.

LangGraph "really aims to minimize the overhead of implementing these patterns." And then the two things it explicitly does not do, which he repeats twice in thirty seconds because it is the whole pitch: it does not abstract prompts, and it does not abstract architecture. What is happening is that LangGraph is "supporting infrastructure underneath any workflow or agent." You write the prompts. You draw the graph. The framework sits below both.

The infrastructure is three things, and he names them in order:

  1. Persistence. This gives you memory. It also gives you human in the loop, "the ability to pause for example while an agent is processing and approve a tool call, pause and review what the agent is doing." He calls this extremely useful in many cases, and it is the item he comes back to in the closing minute.
  2. Streaming. Model APIs stream tokens, he grants, everyone knows that. But when you build these workflows and agents you often want to stream something else: what a particular step output, the output of a tool call. "You often want more flexibility over streaming than just what the LLM calls produce," and LangGraph gives you control over what you emit from any point in a run. He flags this as very useful specifically when putting these things in production.
  3. Deployment. Testing, debugging and deploying. "It is extremely easy to go from any workflow or agent you implement to a deployment."

Then the summary line, stated flatly at 03:06: "LangGraph doesn't abstract prompts and it does not abstract architecture. It's really giving you low level infrastructure that sits underneath any of these workflows or agents." Note what is absent from that pitch. There is no claim that the framework makes the patterns better, shorter or smarter. The claim is that the patterns are yours and the plumbing is theirs.

The building block: the augmented LLM

The post's first figure is the augmented model call, and Lance treats it as the atom every later pattern is assembled from. Models can be augmented with many things. Memory is one, and that is one of the things a framework like LangGraph provides. Models can also interact with tools, "and this is the foundation for building many workflows and agents."

So he builds it from scratch. The notebook setup is three lines of housekeeping: pip install a few packages, set the API key, create the model. The model he runs on through the whole video is Claude, which is visible at 08:42 when he narrates the chain's last step as "Claude creates final joke."

Augmentation one, structured output. Take a schema, in this case a Pydantic model, and bind it to the model. He uses LangChain's with_structured_output method to do it, and then immediately adds the caveat that matters for anyone evaluating the stack: "again, LangGraph does not require you to use LangChain. You can use the raw model APIs, it's completely fine." He runs the cell and the output adheres to the schema he passed.

# a schema, then a model that is contractually obliged to fill it
class SearchQuery(BaseModel):
    search_query: str = Field(None, description="Query that is optimized for web search.")
    justification: str = Field(None, description="Why this query is relevant to the user's request.")

structured_llm = llm.with_structured_output(SearchQuery)
structured_llm.invoke("How does Calcium CT score relate to high cholesterol?")

Augmentation two, tool calling. He defines a function multiply as a tool the model has access to, uses bind_tools to attach it, and invokes the result with an input likely to elicit a call. The tool call comes back: it takes the input and creates the arguments necessary to actually run the function.

def multiply(a: int, b: int) -> int:
    return a * b

llm_with_tools = llm.bind_tools([multiply])
msg = llm_with_tools.invoke("What is 2 times 3?")
msg.tool_calls

Then the sentence that the last third of the video pays off, said casually at 04:38: "remember, when LLMs are creating tool calls they're really just giving you payloads to actually run that tool. You could then run the tool, pass the output of the tool back to the LLM, and really doing that loop gives you an agent." The agent is already fully specified, nine minutes before he builds it. Everything between here and there is the scaffolding you put around this block instead.

LLM call one invoke, nothing more IN Retrieval and memory context the call did not have OUT, SHAPED Structured output a Pydantic schema, bound OUT, THEN BACK IN Tools the model emits a payload, you run it, the result returns Every pattern below is this one block, arranged differently. Nothing after this adds capability. CHAINING []-[]-[] PARALLEL [] [] [] ROUTING []<{3} ORCHESTRATOR []<{n} EVALUATOR []<=>[] AGENT []<->tools
Figure 2. The augmented model call, built in the notebook before any graph exists. Retrieval and memory supply context, a bound schema shapes the output, and tools let the call reach outside itself. The six boxes along the bottom are the rest of the video: same block, different wiring.

Pattern 1: prompt chaining

The intuition. Each model call processes the output of the previous one. That is the whole pattern.

When to reach for it. "When you've attested you can decompose into a few different LLM calls," in Lance's words, meaning when you have confirmed the task actually splits into a handful of steps rather than hoping it does. The post's figure is an input, call one, a gate on the output of call one, then call two, then call three.

The example he builds. A chain that takes a topic from the user, has the model make a joke, checks that the joke has a punchline, and then improves it twice with two subsequent model calls. Three calls, one gate.

Step one: the state container

This is the part he describes as the only real conceptual work. "All I need to do is basically define a container for everything that I want to modify over my workflow." A TypedDict with four keys: topic, which comes from the user; joke, the output of the first call; improved_joke, the output of the second; final_joke, the output of the third.

class State(TypedDict):
    topic: str
    joke: str
    improved_joke: str
    final_joke: str

The state is not bookkeeping. It is the contract between the nodes, and every pattern in the video begins by writing one.

Step two: draw it before you write it

A small aside with more weight than it looks: "when laying out workflows I actually like to draw them out anyway, so this is kind of cool that the blog post has all these laid out." The blog post's figures are not illustrations to him, they are the design step he would have done on paper. And the instruction that follows is the one to keep: "when you draw these out, think about for example each of these calls or steps is just a different function."

So for this workflow there are three functions, generate_joke, improve_joke and polish_joke, and they are all just model calls.

Step three: nodes read state and return state

"What's interesting is with LangGraph, this container or state that I define is passed into every one of these steps, and I can extract whatever I want from it." He reads state["topic"] to get the topic the user wrote into state, and writes back by returning a dict with the key he wants to populate. generate_joke returns {"joke": ...}. improve_joke populates improved_joke. polish_joke populates final_joke.

def generate_joke(state: State):
    msg = llm.invoke(f"Write a short joke about {state['topic']}")
    return {"joke": msg.content}

def improve_joke(state: State):
    msg = llm.invoke(f"Make this joke funnier by adding wordplay: {state['joke']}")
    return {"improved_joke": msg.content}

def polish_joke(state: State):
    msg = llm.invoke(f"Add a surprising twist to this joke: {state['improved_joke']}")
    return {"final_joke": msg.content}

"So what's happening is at each of these steps I'm making LLM calls and I'm populating my state, or this container in the workflow, with the output of each LLM call. That's really all that's going on." A node is a function from state to a partial state update. That sentence describes every node in the video.

Step four: the gate

The gate is not a node. It is a plain function that takes state and returns a string, and the string is what the graph routes on. His check: does the joke contain a question mark or an exclamation point. He is explicit that this is "just some arbitrary criteria," picked because it is cheap to demonstrate, and the function can return anything you want. He returns the two strings Pass and Fail.

def check_punchline(state: State):
    if "?" in state["joke"] or "!" in state["joke"]:
        return "Pass"
    return "Fail"

"Now this can be used as conditional edge or gate in LangGraph." One wrinkle worth noticing, because it is visible on screen and he does not dwell on it: in his wiring, Pass continues into the improvement chain and Fail ends the run. So a joke that already looks like a punchline gets polished twice and a joke that does not gets dropped on the floor, which is the opposite of what a quality gate usually does. With an arbitrary criterion it does not matter for the demo, and the shape of the wiring is the point, but if you lift this code the branch labels are the first thing to think about.

Step five: compile the graph

With a state, three node functions and a gate function in hand, the graph itself is six lines of plumbing.

workflow = StateGraph(State)
workflow.add_node("generate_joke", generate_joke)
workflow.add_node("improve_joke", improve_joke)
workflow.add_node("polish_joke", polish_joke)

workflow.add_edge(START, "generate_joke")
workflow.add_conditional_edges(
    "generate_joke", check_punchline, {"Pass": "improve_joke", "Fail": END}
)
workflow.add_edge("improve_joke", "polish_joke")
workflow.add_edge("polish_joke", END)

chain = workflow.compile()

He narrates it exactly in that order: take in the state, initialize the workflow, add the three steps, "and this is where I define the connectivity of my graph or workflow." Start goes to generate_joke. Then the conditional edge hangs off generate_joke and dispatches on the gate's return value. Then two fixed edges carry the rest. Compile, and "there we go, that's the exact same workflow that they drew out in the blog post, applied to joke generation."

Step six: run it

"Once I have this workflow I can very simply just run chain.invoke. That's a very simple way to invoke any chain, workflow or whatever you have in LangGraph." He passes a topic from the user, and the run does what the graph says: the model creates an initial joke, improves it, and Claude creates the final joke.

THE GRAPH START generate_joke gate check_punchline "?" or "!" in joke Pass improve_joke polish_joke END Fail END THE STATE CONTAINER, PASSED INTO EVERY NODE topic: str written by the user joke: str written by generate_joke improved_joke: str written by improve_joke final_joke: str written by polish_joke A node is a function from state to a partial state update. It reads the keys it needs and returns only the keys it writes. A gate is a function from state to a label. It writes nothing. The graph maps its label to the next node. Solid edges are fixed and always taken. Dashed edges are conditional, and LangGraph draws them dashed in its own render too. Who decides the path: your code, entirely. The gate here is ordinary Python, not a model call. Cost: three sequential model calls instead of one. You trade latency for a simpler instruction at each step.
Figure 3. Prompt chaining compiled, with the state container drawn underneath and every node mapped to the single key it writes. This is the template for every other pattern in the video: one state object, nodes that are functions, and edges that are either always taken or chosen by a label.

Pattern 2: parallelization

When to reach for it. Two cases, and he gives a real one from his own work for each. First, when you want multiple perspectives on a single task: "I've used this quite a bit with things like multi query RAG. If I have a question, I want to fan it out into like three or four different sub questions." Second, when independent tasks can be performed with different prompts or different models. "Lots of cases where you want to just parallelize tasks in workflows."

The example. Take a topic and create a joke, a story and a poem, all at once, then combine them.

The state. Same move as before, "just a container for everything I'm going to modify in my workflow": a topic from the user, then joke, story, poem, and combined_output to hold the aggregate.

class State(TypedDict):
    topic: str
    joke: str
    story: str
    poem: str
    combined_output: str

The nodes. Three model calls and one aggregation. call_llm_1 writes a joke about the topic, call_llm_2 a story, call_llm_3 a poem, and aggregator reads all three out of state and joins them into a single string. "Super simple," he says, and it is: four functions, none longer than three lines.

The graph. This is the shortest build in the video, because the parallelism is not a feature you switch on. It is a consequence of the edges you drew.

parallel_builder = StateGraph(State)
parallel_builder.add_node("call_llm_1", call_llm_1)
parallel_builder.add_node("call_llm_2", call_llm_2)
parallel_builder.add_node("call_llm_3", call_llm_3)
parallel_builder.add_node("aggregator", aggregator)

parallel_builder.add_edge(START, "call_llm_1")
parallel_builder.add_edge(START, "call_llm_2")
parallel_builder.add_edge(START, "call_llm_3")
parallel_builder.add_edge("call_llm_1", "aggregator")
parallel_builder.add_edge("call_llm_2", "aggregator")
parallel_builder.add_edge("call_llm_3", "aggregator")
parallel_builder.add_edge("aggregator", END)

parallel_workflow = parallel_builder.compile()

"In this case I go from start to LLM call one, two, three, and they all connect to the aggregator on the end." Three edges out of START means three nodes with no dependency between them, so they run together. Three edges into aggregator means it waits for all three. He displays the graph image, calls it "a nice visualization of your workflow," runs it, and gets a story, a joke and a poem created in parallel.

What it buys you. Wall clock time in the sectioning case. In the voting case, which the post also covers, running the same task several times and aggregating buys reliability on judgment calls where a single sample is noisy.

Pattern 3: routing

When to reach for it. "I've used routing a lot and it's extremely useful in a bunch of use cases." The trigger is wanting to go to exactly one of several steps, and wanting the model to make that choice. His concrete example from production work: "I've used routing a lot in working with retrieval, in taking a question and routing it to different retrieval systems."

The example. Take an input and route it to joke, story or poem generation based on what the user actually asked for.

The trick he wants you to steal. "Now I'm going to show a very nice trick here for routing, I often like to do this." Give the routing model a structured output, so it is guaranteed to produce one of the three strings as a structured object rather than prose you have to parse. The schema is a Pydantic model with a single field typed as a Literal of the three options.

class Route(BaseModel):
    step: Literal["poem", "story", "joke"] = Field(None, description="The next step in the routing process")

router = llm.with_structured_output(Route)

"That guarantees the LLM is going to produce, in this particular case, the strings poem, story or joke as a structured object." The guarantee is the point. Without it the router node becomes string matching against a sentence, and that is where routers break.

The state. An input from the user, a decision from the router, and an output saved as the final result.

The nodes. Three model calls that write a story, a joke or a poem, plus the router node. The router takes the user input, decides which of the three the content calls for, returns the Pydantic object, and then Lance does the one extra step people forget: he extracts .step off the returned model and writes it into state as decision.

def llm_call_router(state: State):
    decision = router.invoke([
        SystemMessage(content="Route the input to story, joke, or poem based on the user's request."),
        HumanMessage(content=state["input"]),
    ])
    return {"decision": decision.step}

def route_decision(state: State):
    if state["decision"] == "story":
        return "llm_call_1"
    elif state["decision"] == "joke":
        return "llm_call_2"
    elif state["decision"] == "poem":
        return "llm_call_3"

"So then I have my decision in state, and now this is just going to be a conditional edge that'll look at the decision and determine what node to go to. Super simple. So basically if the decision is story go to node one, joke two, poem three. That's it."

The honesty about the toy. He breaks frame deliberately here, because the example is degenerate: "you'll see in this toy example each of these steps are doing the same thing, but in real world examples the router will send to different steps that will have different logic, different LLM calls. This is just a toy example showing you how to hook up that logic between for example a router step, a structured output, and a decision about where to go next."

The rendering detail. When he visualizes the compiled graph he points at something concrete: "what's kind of nice in LangGraph, when we visualize this, this dotted line means a conditional edge. So it's going to go to only one of those three paths whenever this runs, and of course that's going to be based upon the decision of the router." A dotted edge in a LangGraph render is a promise about runtime, not a stylistic choice, and it is how you read someone else's graph at a glance.

The run. He adds a print statement to each of the three nodes so the output tells you which node was visited, runs it, confirms the joke node fired, and gets the joke back.

PARALLELIZATION, all three run START call_llm_1 joke call_llm_2 story call_llm_3 poem aggregator waits for all three ROUTING, exactly one runs START llm_call_router Route.step -> decision structured output, so the label is guaranteed legal llm_call_1 story llm_call_2 joke llm_call_3 poem Same silhouette, opposite semantics. The edge style is the only thing that tells you which one you are reading. Solid fan-out from START: no dependency between the branches, so LangGraph runs them together and the join waits for all of them. Dashed fan-out from a node: a conditional edge, so exactly one branch is taken, chosen by a label a model produced. Parallelization buys wall clock time, or reliability when you run the same task repeatedly and vote. Your code still owns the path. Routing buys specialization: each branch gets its own prompt, its own logic, its own tools, instead of one prompt covering everything.
Figure 4. Parallelization and routing, drawn to the same scale because they are easy to confuse on a whiteboard. In LangGraph's own rendering a dotted edge is a guarantee that exactly one path is taken at runtime, which Lance points out on screen as the way to read any compiled graph.

Pattern 4: orchestrator and workers

This is the one he spends the longest on, and the one with a genuinely new API in it.

What it is. "A case where you want an LLM to break down a task into a set of subtasks, delegate each subtask to an independent worker, and then synthesize the results."

How it differs from parallelization. This is the sentence to hold onto: "it's kind of like parallelization, except the key difference is this worker assignment you don't know ahead of time. So you're having an LLM reason about something and then create a bunch of workers based upon its reasoning." And then the structural consequence: "so again in this case the LLM is kind of gating or creating the control flow, just like in the case of routing." Parallelization has a fan-out width you typed into your graph. Here the width is an output of a model call.

The example. Report writing, which he says he has used a lot. "Maybe you've played with deep research. An LLM reasons about the plan for the report and dynamically generates a bunch of report sections, and then goes and does research on all of them. Classic example of an orchestrator worker type workflow." LangChain later shipped Open Deep Research as an open source build of exactly this shape, if you want to read a production sized version of the pattern.

The planner is structured output again

"For this, the trick is I'm also going to use structured outputs." Two Pydantic models: a Section with a name and a description, and a Sections wrapper holding a list of them. Bind the list schema to the model and that bound object is the planner.

class Section(BaseModel):
    name: str = Field(description="Name for this section of the report.")
    description: str = Field(description="Brief overview of the main topics and concepts to be covered in this section.")

class Sections(BaseModel):
    sections: List[Section] = Field(description="Sections of the report.")

planner = llm.with_structured_output(Sections)

"So what's cool here is the planner is going to take an input, reflect on it, and produce a list of sections based upon its reflection. So this is dynamic, I don't know how many sections it'll create a priori. That's why this is a very good orchestrator worker use case." The test for whether you need this pattern is exactly that: if you can count the branches before the run, you do not need it.

Two states, not one

This is the part he flags as the real design wrinkle, and it is the only place in the video where a pattern needs more than one state object.

The orchestrator's graph state has four keys: topic from the user, sections (the planner's list), completed_sections (which all the workers write into), and final_report.

The workers get their own state. "This is where things get interesting with orchestrator worker workflows in LangGraph. The way we often like to do it is, for the workers, give them their own state." His reason: "each of those workers, you want to handle independent inputs, and they're kind of all self-contained objects. Think about them as their own little buckets in which work's being done, and different work's being done in each one, but they're all writing out to the same output."

class State(TypedDict):
    topic: str
    sections: list[Section]
    completed_sections: Annotated[list, operator.add]
    final_report: str

class WorkerState(TypedDict):
    section: Section
    completed_sections: Annotated[list, operator.add]

"And this is why I include this completed_sections key in the worker state and in the graph state." Then the mechanism: "what's interesting in LangGraph is when you have overlapping keys, and you write to for example this completed_sections key in each worker, the outer state will also have that update reflected."

The piece that makes the parallel writes safe is the annotation. "All the workers are going to write to this completed_sections key in parallel, and we structure this key with an annotation that allows for the addition of new elements." That is a reducer: Annotated[list, operator.add] tells LangGraph that two concurrent writes to this key should be concatenated rather than one overwriting the other. Without it, four workers finishing at once would leave you with one section. "So that's really all we need to do."

The three node functions

The orchestrator is the planner, invoked with the topic and told to create a plan. The worker takes WorkerState, is handed a section name and description, writes that one section, and returns it under completed_sections. The synthesizer reads completed_sections out of state, joins them into a string, and writes final_report.

def orchestrator(state: State):
    report_sections = planner.invoke([
        SystemMessage(content="Generate a plan for the report."),
        HumanMessage(content=f"Here is the report topic: {state['topic']}"),
    ])
    return {"sections": report_sections.sections}

def llm_call(state: WorkerState):
    section = llm.invoke([
        SystemMessage(content="Write a report section following the provided name and description."),
        HumanMessage(content=f"Here is the section name: {state['section'].name} and description: {state['section'].description}"),
    ])
    return {"completed_sections": [section.content]}

def synthesizer(state: State):
    completed_report_sections = "\n\n---\n\n".join(state["completed_sections"])
    return {"final_report": completed_report_sections}

"This completed_sections is a state key that all the workers can write to in parallel, that's the key point, and all those sections are going to be accumulated in that completed_sections key. Then I'm going to have a synthesizer that's going to read out the completed sections and just write them all out as a string. That's it."

The one genuinely new API: Send

"Now this is the only thing that's new and a little bit special in orchestrator worker style workflows. Because this is so common, we have a special API called send in LangGraph that allows you to basically spawn these workers dynamically."

The mechanism is small. The planner already wrote the sections into state. You write a function that iterates them and returns a list of Send objects, one per section, each naming the node to run and the state to initialize it with. You hang that function off the orchestrator as a conditional edge.

def assign_workers(state: State):
    return [Send("llm_call", {"section": s}) for s in state["sections"]]

"For each section in sections, send to llm_call, that's my worker, and basically initialize the state as section when I do that. So then that llm_call receives worker state, and it's receiving from state the section name and description, and then it goes and writes that section. It writes its output to completed_sections, which my orchestrator has access to, and then the synthesizer basically just grabs completed sections and combines them. That's it, you're done, that's all you need to do."

orchestrator_worker_builder = StateGraph(State)
orchestrator_worker_builder.add_node("orchestrator", orchestrator)
orchestrator_worker_builder.add_node("llm_call", llm_call)
orchestrator_worker_builder.add_node("synthesizer", synthesizer)

orchestrator_worker_builder.add_edge(START, "orchestrator")
orchestrator_worker_builder.add_conditional_edges("orchestrator", assign_workers, ["llm_call"])
orchestrator_worker_builder.add_edge("llm_call", "synthesizer")
orchestrator_worker_builder.add_edge("synthesizer", END)

orchestrator_worker = orchestrator_worker_builder.compile()

Reading the rendered graph

Another concrete note on the visualization, and this one is useful: "this dotted line just shows you that you use the send API to spawn a whole bunch of llm_call workers. You don't know how many, a heap ton, that's why it doesn't draw out each one specifically. It's going to be dynamically determined based upon the orchestrator plan. That's the key characteristic of these orchestrator worker workflows: you don't know a priori how many workers you need, the LLM will determine that on the fly."

So in a LangGraph render, a single dotted edge into one worker node is not a single worker. It is an unknown number of them, and the picture cannot tell you more than that.

The run

He asks for a report on LLM scaling laws. "What's kind of cool is this runs fairly quickly." The planner generates the report sections, he opens final_report, and the sections are concatenated in plan order: "you can see I get this rich introduction, and then I get, you know, fundamentals, scaling relationships, following the plan."

Then the caveat, which is the kind of thing that separates a demo from a claim: "now in reality when I do this I have much more detailed prompts than I showed here. This is really showing you the workflow rather than the particulars of how I build a high quality report writer. In fact I have separate videos on report writing you could check out, but this is just showing you how to set up an orchestrator worker style of workflow." The graph shape is the deliverable. The prompt quality is a different, longer problem.

START orchestrator planner, bound to Sections schema returns list[Section] length unknown until it runs Send assign_workers llm_call WorkerState one worker per Send, and the planner decides how many, not you all of them write to the same key, in parallel completed_sections Annotated[list, operator.add] the reducer concatenates the parallel writes instead of one of them overwriting the rest synthesizer joins -> final_report END THE DISTINCTION THAT MATTERS Parallelization You wrote three edges out of START, so there are three branches. Forever. The width of the fan-out is a fact about your source code. add_edge(START, "call_llm_1") x3 Orchestrator and workers A model call returns a list, and the length of that list is the width of the fan-out. The model is now writing your control flow. [Send("llm_call", ...) for s in sections]
Figure 5. The orchestrator and workers graph, with the two details that actually make it work: the Send API turning a planner's list into a dynamic number of worker invocations, and the operator.add reducer on completed_sections that lets those workers write to one key at once. The panel below is the test for whether you need this pattern at all.

Pattern 5: evaluator and optimizer

What it is. One model generates a response, another grades it and gives feedback, in a loop. He introduces it as the sibling of the previous pattern: "if we just talked about the orchestrator worker workflow, let's talk about something that's a bit related, the evaluator optimizer workflow. In both these cases you're going to have LLMs directing the control flow through predefined code paths."

Where he has used it. Grading RAG outputs. "I've used this quite a bit, for example grading responses from a RAG system for hallucinations or for factual accuracy. I've used these kind of like evaluator gates, and if for example there's hallucination in the output that isn't grounded by the documents, I send it back and have it regenerate a response." That is the canonical use: the quality bar is checkable even when it is not reliably hittable on the first try.

The structured outputs aside, which is the best thirty seconds in the video

Before building it he stops to generalize, because this is the fourth pattern in a row that leans on the same primitive. "Here I'm going to use structured outputs again. You see there's kind of a trend here. Structured outputs is like, kind of, all you need, in quotes."

Then the argument: "it's extremely convenient because you can really build all these workflows just off structured outputs. You don't have to use tool calling, for example. You can use routers to basically conditionally determine where to go next, you can have then nodes that for example just call tools depending on the result of the router itself. So really, just structured outputs is a nice way you can build a lot of complex workflows."

That is a real engineering position and worth stating plainly: for workflows, a model that reliably fills a schema is enough. Tool calling is a capability you need for agents, not a prerequisite for scaffolding. Four of the five workflow patterns in this video route on a schema field, not a tool call.

The grader schema

class Feedback(BaseModel):
    grade: Literal["funny", "not funny"] = Field(description="Decide if the joke is funny or not.")
    feedback: str = Field(description="If the joke is not funny, provide feedback on how to improve it.")

evaluator = llm.with_structured_output(Feedback)

"That's going to be my my grader model, that's going to be grade and feedback. Decide if the joke is funny or not, in my case, and if it is not funny give some feedback."

The state and the loop

The state carries a topic, the joke, the feedback string and the funny_or_not grade. He describes the cycle in one breath: "I'm going to take a joke based upon a topic from a user, I'll generate the joke, and then I'll grade it, give it feedback, determine if it's funny or not. Based on the feedback I'll go back, regenerate a new joke. That's it, nice and simple."

The generator has the one piece of logic that makes the loop work, and he flags it himself: "you'll see I do something kind of interesting here. I'm going to check if there's feedback in the state. Now there might be, because I've basically looped back to this node if I determine that the joke is not good. So there may be feedback in the state. If there's feedback, I include it in my prompt, otherwise I just say write a joke about the topic."

def llm_call_generator(state: State):
    if state.get("feedback"):
        msg = llm.invoke(f"Write a joke about {state['topic']} but take into account the feedback: {state['feedback']}")
    else:
        msg = llm.invoke(f"Write a joke about {state['topic']}")
    return {"joke": msg.content}

def llm_call_evaluator(state: State):
    grade = evaluator.invoke(f"Grade the joke {state['joke']}")
    return {"funny_or_not": grade.grade, "feedback": grade.feedback}

def route_joke(state: State):
    if state["funny_or_not"] == "funny":
        return "Accepted"
    elif state["funny_or_not"] == "not funny":
        return "Rejected + Feedback"

One node, two prompts, branching on whether a state key is populated. That is the entire difference between the first pass and every pass after it, and it is why the loop does not just regenerate the same joke forever.

The conditional edge that closes the cycle

"Then I have a conditional edge that'll look at state funny or not, so again that's like kind of my grade, and if it's funny route to accepted, if it's not funny route to rejected and feedback."

optimizer_builder = StateGraph(State)
optimizer_builder.add_node("llm_call_generator", llm_call_generator)
optimizer_builder.add_node("llm_call_evaluator", llm_call_evaluator)

optimizer_builder.add_edge(START, "llm_call_generator")
optimizer_builder.add_edge("llm_call_generator", "llm_call_evaluator")
optimizer_builder.add_conditional_edges(
    "llm_call_evaluator",
    route_joke,
    {"Accepted": END, "Rejected + Feedback": "llm_call_generator"},
)

optimizer_workflow = optimizer_builder.compile()

"If accepted we go to end, if it is rejected and feedback we go back to llm_call_generator. So you can see, when you set the conditional edge, this is how you can basically route from the output of your logic, your edge logic, to the next node to go to. That's it, it's extremely simple, and you can see you get a nice visualization of it."

This is the first graph in the video with a cycle in it, and the cycle is not a special construct. It is a conditional edge whose target happens to be a node that already ran. That is the thing a graph gives you that a chain does not.

The run

Input: cats. He gets the joke out, then inspects the state to see what the grader said. state["feedback"]: "in this case it seems to like it." state["funny_or_not"]: funny. "So the grader liked the joke, we went ahead and returned it to the user, we ended. So we can see that the grader was initiated and decided to like the joke, so it passes and we finish." One pass, no revision, which is a slightly anticlimactic demo of a loop and he does not pretend otherwise.

Rejected + Feedback, route_joke sends it back to the generator START llm_call_generator if state.get("feedback"): fold it into the prompt first pass: topic only later passes: topic plus critique joke llm_call_evaluator bound to Feedback schema grade + feedback Literal["funny", "not funny"] Accepted END THE CYCLE Nothing new was added to make the loop. It is a conditional edge whose target is a node that already ran, which is exactly what a graph lets you write and a linear chain does not. The generator checks for a state key rather than taking a flag, so it never needs to know which pass it is on. Worth noting before you ship it: nothing in this graph caps the number of revisions. The loop exits only when the grader says funny.
Figure 6. The evaluator and optimizer cycle, the first graph in the video that is not a directed acyclic graph. The generator's one conditional on the presence of a feedback key is what distinguishes the first attempt from every revision after it.

Removing the scaffolding

At 23:32 he draws the line under the workflow half of the video and restates what has actually been happening: "we talked about a number of different workflows that use LLMs within some kind of reasoning scaffolding, and in the case of orchestrator worker, evaluator optimizer, routing, you actually do let the LLM make decisions to route the control flow through that scaffolding."

That is the middle band of Figure 1, named explicitly, with its three members listed. Three of the five workflows already hand control flow decisions to a model. What makes them workflows is that every path the model can choose is a path he typed.

"Now let's remove the scaffolding and let's talk about agents. With agents you're simply allowing an LLM to perform actions in the form of tool calls, and directly receive the output or feedback from those actions. And so in the workflow case we talked about, there are always kind of these predefined code paths that we had an LLM follow and route through. In the case of an agent, we've removed those."

When do you actually need one

"This is kind of the big question." His answer has two halves, and the second half is the more useful one.

The case for. "You see agents being used in cases where you really have open ended problems that you cannot easily capture in a workflow. For example, you want the LLM to utilize different tools in a pattern that you just cannot predict a priori, so it's not easy to lay it out in the workflow." The existence proof he cites is SWE-bench, "a benchmark for software engineering in which Anthropic actually used an agent architecture, just shown like this, and achieves very strong performance." That is a real result: Anthropic's write-up on SWE-bench Verified describes exactly this shape, a model in a loop with bash and file editing tools and very little scaffolding. "So we know for certain open ended tasks, agents are appropriate."

The caveat he puts right next to it. "I do want to caveat, in the event that LLMs get extremely proficient at tool calling, it's also possible that a lot of the scaffolding that we talked about with various workflows is unnecessary. Today, if you know roughly the sequence tools need to be initiated, it's often better just capture it in a workflow, in terms of reliability, than just give it to an agent and let the agent hopefully call that correct sequence."

Read those two paragraphs together and you have the decision rule the whole video is built around. The question is not how smart the model is. It is whether you know the sequence. If you know it, write it down, because a written sequence runs the same way every time and a hoped-for sequence does not.

Pattern 6: the agent loop, built from scratch

"Now let's just set up an LLM with tools." Three tools, all arithmetic: multiply, add and divide. "Nice and simple."

def multiply(a: int, b: int) -> int:
    return a * b

def add(a: int, b: int) -> int:
    return a + b

def divide(a: int, b: int) -> float:
    return a / b

tools = [add, multiply, divide]
tools_by_name = {tool.name: tool for tool in tools}
llm_with_tools = llm.bind_tools(tools)

The state is one key

The whole agent runs on a single state key. "My state has a single key, in this particular case messages, which is going to accumulate what I pass as the user, what the LLM produces, and so forth." That is MessagesState, a prebuilt state whose one key is a message list with an append reducer on it, the same mechanism as completed_sections in the orchestrator but applied to a conversation instead of report fragments.

Two nodes

llm_call. The model with tools bound, given a system message: "you're a helpful assistant tasked with performing arithmetic." Its output is appended to messages.

tool_node. "It's going to look at the state, look at the last message, determine if it's a tool call. If it is, it'll go ahead and just call that tool. That's it, nice and simple. And it's going to return that to state."

def llm_call(state: MessagesState):
    return {"messages": [llm_with_tools.invoke(
        [SystemMessage(content="You are a helpful assistant tasked with performing arithmetic on a set of inputs.")]
        + state["messages"]
    )]}

def tool_node(state: dict):
    result = []
    for tool_call in state["messages"][-1].tool_calls:
        tool = tools_by_name[tool_call["name"]]
        observation = tool.invoke(tool_call["args"])
        result.append(ToolMessage(content=observation, tool_call_id=tool_call["id"]))
    return {"messages": result}

Environmental feedback, defined

This is the clearest definition of the term anywhere in the video, and he gives it while tracing the message list: "I'm going to have a sequence of human input, model, agent in this case decides to call a tool, this tool_node looks and sees, oh, the LLM decided to call a tool. It actually runs that tool call, that then is written to state messages as a tool message. This is that environmental feedback thing you hear about. So you talk about agents, they can perform actions, that's the tool call done up here, fine. They also can receive feedback from the environment and act on it. That is the output of this tool node. This tool node is basically the environmental feedback saying here's the output of the tool call. The LLM then will get that, decide what to do next. That's it, that's an agent."

The term stops being mystical once you see it is a ToolMessage appended to a list.

One conditional edge

"The only other thing I need is this conditional edge that basically says, was the last message a tool call? If so I'm going to go ahead and route to the tool node, and if not I will end. You can modify this in different ways, but a lot of times people basically just say allow the agent to continue making tool calls until it decides it doesn't need one anymore, and then you're done."

def should_continue(state: MessagesState) -> Literal["tool_node", END]:
    last_message = state["messages"][-1]
    if last_message.tool_calls:
        return "tool_node"
    return END

agent_builder = StateGraph(MessagesState)
agent_builder.add_node("llm_call", llm_call)
agent_builder.add_node("tool_node", tool_node)

agent_builder.add_edge(START, "llm_call")
agent_builder.add_conditional_edges("llm_call", should_continue, ["tool_node", END])
agent_builder.add_edge("tool_node", "llm_call")

agent = agent_builder.compile()

Count the pieces. One state key, two nodes, one fixed edge back from the tool node to the model, one conditional edge out of the model. That is the whole agent, and the loop exists because tool_node points back at llm_call.

"So there we are, I mean that's our agent loop, and that's kind of why agents are elegant. They're extremely simple in this kind of formulation. All it is, is basically an LLM initially with a bunch of tools, and it's like a tool node that will run the tool for you and return that environmental feedback to the LLM, and let the LLM keep spinning until it decides it doesn't need a tool call anymore, then you're done. That's really it."

Who owns the actions

A small but load bearing point about the division of labor: "now again, this tool call thing, think about that as actions, so this LLM can perform actions. The actions are determined by the user, so you can pass in any set of tools to this LLM, it decides to perform those actions. Now you need some system that actually does those actions, so that's what we create with this tool node. And this just loops until the LLM says I don't need a tool call anymore, and then that conditional edge says okay, I can just end."

You choose the action space. The model chooses the sequence. Your code executes. That split is the entire security and reliability surface of an agent, stated in one paragraph.

The run

"I'm going to basically tell this agent: add three and four, then take the output and multiply by four."

And then the self aware bit, because the demo is trivially solvable without tools: "look, this is extremely simple and an LLM can just do this, I totally get it. This is more showing you the principle: setting up an agent with these tools for addition and multiplication, and testing whether or not it can correctly perform these tool calls in sequence, and looking at the flow of messages."

The flow of messages is the deliverable, and he walks it: "you have an input from the human, my instructions. The LLM looks at that and says, okay, I need to make a tool call, makes a tool call. The tool node executes the tool call, returns environmental feedback to my LLM as the result of the tool call, which is seven. My agent thinks about it, makes another tool call, responds with twenty eight. Agent thinks about it, final result is twenty eight, here's how we got there, no more tool calls needed, done."

Then the line the video is remembered for, at 28:35: "for all the hype agents, that's all it is: extremely simple, is tool calling in a loop."

THE GRAPH: TWO NODES, ONE LOOP appends a ToolMessage: this is the environmental feedback START llm_call model + bound tools should_continue tool_node runs the payload no tool_calls on the last message END state: MessagesState one key, messages, append reducer No step cap, no cost ceiling. The loop exits when the model stops asking for tools. THE ACTUAL RUN: "ADD 3 AND 4, THEN TAKE THE OUTPUT AND MULTIPLY BY 4" 1 HumanMessage Add 3 and 4, then take the output and multiply by 4 2 AIMessage tool_calls: add(a=3, b=4) action 3 ToolMessage 7 environmental feedback 4 AIMessage tool_calls: multiply(a=7, b=4) action, chosen from the feedback 5 ToolMessage 28 environmental feedback 6 AIMessage Final result is 28, and here is how we got there. No tool calls. no tool_calls, so END
Figure 7. The agent built from scratch, and the six message list it produces on the arithmetic run. Note that nothing decided in advance that there would be two tool calls. The loop ran twice because message 6 was the first AI message without a tool call on it, and that is the only termination condition in the graph.

So why not use agents for everything

Having just called the agent elegant, he spends the next minute arguing against reaching for it, and this is the most quotable stretch of the video because it is a practitioner reporting what he actually ships.

"Now again, why don't we just do this? Because look, a lot of problems can be solved with workflows, which are a bit simpler. Agents haven't been particularly reliable to date, particularly with large numbers of tools or complex trajectories of tool calls, and so a lot of people actually in production prefer workflows. I've done a lot more workflows than agents, to be honest, over the last year or two."

Two specific failure conditions are named there, and they are worth separating because they fail differently. Large numbers of tools is a selection problem: the model has to pick correctly from a wide menu on every turn. Complex trajectories of tool calls is a compounding problem: a long sequence multiplies per step error rates until the run ends somewhere unrelated to the goal. A workflow removes the first by constraining the menu per node and removes the second by fixing the order.

Then the forecast, hedged precisely: "I completely acknowledge, in the event that you have very capable reasoning models that can perform low latency high quality tool calling, and we have more confidence that they can perform reliably in production, I think you will see the movement to this classic style of very simple tool calling agent in production. But the game in the field today is a little bit more like, people are saying, well, I want to put workflows into production because they're a little bit more trustworthy. Again, I have some scaffolding around the core LLM calls."

Note what is doing the work in that conditional. Not capability, but confidence. The blocker is not that models cannot call tools, it is that teams cannot yet predict how a given agent will behave on inputs they have not seen, and scaffolding is how you buy predictability back.

The prebuilt shortcut

At 29:35 he opens the documentation and shows the thing he deliberately did not use: "I do also want to show, this is looking at the documentation, that we do have a prebuilt method called create_react_agent that basically wraps what we just built from scratch right here, for convenience."

And he refuses to have an opinion about which you should use: "now this depends. If you want to build it from scratch yourself just as we did, that is completely fine. If you want to use the prebuilt method, that's fine as well, but it's available to you."

He also commits to sharing the notebook: "I will be sharing the link to this tutorial page which has everything we just went through, so all the different workflows we talked through are all in this tutorial." That page is Workflows and agents in the LangGraph docs, and every snippet above is in it, in Python and TypeScript, which is the single most useful artifact attached to this video.

The framework argument, restated with the code behind it

The last ninety seconds return to the three claims from minute two, now that you have watched eight graphs get compiled. "If I go back to the why LangGraph story, when you compile those workflows or agents in LangGraph, you're getting a persistence layer for free."

Persistence. "Which gives you short and long term memory, and that also gives you the ability to stop, perform interruptions, review and continue, IE human in the loop." Three capabilities from one mechanism: memory across turns, resumption after a stop, and a review gate in the middle of a run.

Streaming. "It gives you a whole bunch of streaming capacities. You can stream independent values from your state at any point in time, you can stream of course tokens out of LLM calls." The first half is the one that is hard to get otherwise. Streaming tokens is a model API feature; streaming an arbitrary state key as a node writes it is a framework feature, and it is what makes a multi minute orchestrator run feel alive rather than hung.

Deployment. "We have a very nice and easy onramp for testing, debugging and deploying any of these, so you could take any of those workflows we just built and deploy them in like five minutes. So very, very quickly, that's really what you're getting with the framework."

He closes on the direction of travel, which is also an admission that the overhead is real: "and you can also see with LangGraph it's pretty easy to lay them out, and we're actually working on making it even easier. So we're trying to reduce the overhead, so you can lay it out almost as if you're writing Python, you're not even thinking about the framework, but you're getting these benefits kind of for free when you do use a framework. That's kind of the big idea here."

Then: "hopefully this was a helpful overview to present how to lay these various workflows or agents out using LangGraph, and what benefits you may get from LangGraph as a consideration. So thanks very much, feel free to leave any comments below."

Every pattern, as a reference table

PatternWhat it isWho picks the pathLangGraph primitiveReach for it whenWrong choice when
Augmented LLMOne model call with structured output, tools, retrieval or memory attachedyour codewith_structured_output, bind_toolsAlways. It is the atom every pattern below is built fromNever, but on its own it has no control flow at all
Prompt chainingA fixed sequence of calls, each consuming the last one's output, with an optional gate partwayyour codeadd_edge in a line, plus one add_conditional_edges for the gateYou have confirmed the task decomposes into a few steps, and each step gets a simpler instructionThe steps are independent. You are paying serial latency for nothing
ParallelizationSeveral independent calls at once, joined by an aggregator. Sectioning or votingyour codeMultiple add_edge calls out of START, all converging on one nodeYou want several perspectives on one task, or independent subtasks with different prompts or modelsA later branch needs an earlier branch's output. The join waits for everything
RoutingClassify the input, then run exactly one specialized brancha model call, from a fixed menuadd_conditional_edges on a Literal field from structured outputOne prompt is trying to cover several different jobs and doing all of them adequatelyYou can classify with a regular expression. The model call is pure cost
Orchestrator and workersA planner decides the subtasks, workers run them in parallel, a synthesizer combinesa model call, menu and widthSend from a conditional edge, plus Annotated[list, operator.add] on the shared keyYou genuinely cannot count the branches before the run. Report writing, deep researchYou can count them. Use parallelization and keep the width in your source code
Evaluator and optimizerA generator and a critic in a cycle, revising until the grade passesa model call, loop or exitA conditional edge pointing back at a node that already ranQuality is easier to check than to achieve first try. Hallucination grading, style constraintsThere is no checkable criterion, or the check costs as much as the generation. Also uncapped by default
AgentA model with tools in a loop, choosing each next action from environmental feedbackthe model, unboundedMessagesState, a tool node, and one conditional edge on last_message.tool_callsThe sequence of tools genuinely cannot be predicted. Open ended coding, SWE-bench style tasksYou roughly know the sequence. Write it down instead, for reliability. Also unreliable with many tools or long trajectories

Key takeaways

Chapters

Notable quotes

Think about a workflow as some kind of scaffolding of predefined code paths around LLM calls. Lance Martin, the definition the whole video runs on, 0:20

Agents remove the scaffolding, so you're basically letting an LLM direct its own actions. Lance Martin, the other half of the definition, 1:02

Workflows have some scaffolding around them, whereas an agent is unbounded. Lance Martin, after granting that workflows call tools too, 1:25

Implementing these patterns doesn't require framework, right? Sometimes this can be done in a few lines of code. I completely understand that. Lance Martin, opening the case for a framework by conceding the case against, 1:40

LangGraph doesn't abstract prompts and it does not abstract architecture. It's really giving you low level infrastructure that sits underneath any of these workflows or agents. Lance Martin, the pitch, stated twice in thirty seconds, 3:06

Remember, when LLMs are creating tool calls, they're really just giving you payloads to actually run that tool. You could then run the tool, pass the output of the tool back to the LLM, and really doing that loop gives you an agent. Lance Martin, specifying the agent nineteen minutes before building it, 4:40

All I need to do is basically define a container for everything that I want to modify over my workflow. Lance Martin, on state, the first step of every build in the video, 5:40

When laying out workflows I actually like to draw them out anyway, so this is kind of cool that the blog post has all these laid out. Lance Martin, on why the post's figures are a design step and not decoration, 5:58

You don't know a priori how many workers you need. The LLM will determine that on the fly. Lance Martin, on the one test for whether you need orchestrator and workers, 18:53

Structured outputs is like, kind of, all you need, in quotes. Lance Martin, after using it in four patterns running, 19:55

This is that environmental feedback thing you hear about. Lance Martin, pointing at a ToolMessage in a list, 26:10

That's it. That's an agent. Lance Martin, having written two node functions, 26:35

For all the hype agents, that's all it is: extremely simple, is tool calling in a loop. Lance Martin, after the arithmetic run returns 28, 28:50

Agents haven't been particularly reliable to date, particularly with large numbers of tools or complex trajectories of tool calls, and so a lot of people actually in production prefer workflows. Lance Martin, the honest state of the field in January 2025, 29:05

I've done a lot more workflows than agents, to be honest, over the last year or two. Lance Martin, reporting on his own work, 29:15

If you know roughly the sequence tools need to be initiated, it's often better just capture it in a workflow, in terms of reliability, than just give it to an agent and let the agent hopefully call that correct sequence. Lance Martin, the decision rule, 24:45

You could take any of those workflows we just built and deploy them in like five minutes. Lance Martin, closing on the third framework claim, 31:08

What the captions mangled

The auto captions on this video are rough on proper nouns, so a single correction pass rather than quiet edits throughout. Everything below is corrected in the quotes and prose above.

Resources mentioned

The source material

The framework and the primitives he uses

Things he references in passing

An honest footnote

Two things to hold alongside this video, neither of which undercuts it.

The forecast is the part to watch. Lance's own hedge at 24:33 and again at 29:05 is that the workflow preference is a statement about today, not about the shape of the problem: if models become reliable enough at low latency tool calling, much of the scaffolding becomes unnecessary and the simple tool calling agent moves into production. He recorded that in January 2025. That is a falsifiable prediction about a moving target, and it is the lens to read the video through rather than a timeless law. The patterns themselves do not expire, because the question they answer, do you know the sequence or not, does not depend on model capability. The default answer does.

The vendor position is stated, not hidden. This is a LangChain engineer making a case for LangGraph, and the case is unusually narrow: persistence, streaming, deployment, with prompts and architecture left to you. He opens by conceding that the patterns need no framework at all. Hold him to the narrow version. If you do not need checkpointing, human interrupts, cycles or streamed intermediate state, the five workflow patterns really are a few functions and some control flow, and the honest reading of his own argument is that you should write them that way.

A practical note on the code itself, since this page is meant to be usable as a reference: none of the eight builds has a step cap, a cost ceiling or a retry limit. The evaluator loop exits only when the grader says funny, and the agent loop exits only when the model stops asking for tools. That is correct for a teaching notebook and not correct for anything you point at production, and the bounding is on you because, as he says, the framework does not abstract architecture.

Where this sits in the LLM Learning track

This is the last of the four videos in Part 4 of the LLM Learning track, the stage about shipping something, and it closes the building half of the curriculum.

It reads directly out of the two videos before it. A Hackers' Guide to Language Models walks the practical ladder up to function calling and retrieval, which is the capability this video assumes on page one: every pattern here is built on a model that can fill a schema and emit a tool payload. How to Construct Domain Specific LLM Evaluation Systems is the real prerequisite, and the order matters more here than anywhere else in the track. Every pattern on this page multiplies the number of ways a system can fail, and a cycle or a dynamic fan-out multiplies it without bound. Lance's own reason for preferring workflows, that agents have not been reliable with long tool call trajectories, is a claim you can only make about your own system if you are already looking at traces and scoring runs. Read the evals page first, then read this one, and the pairing gives you the shape of a real product: a graph you can draw, an evaluation loop that ratchets, and an agent only where the sequence genuinely cannot be written down.

It also sets up Part 5. Andrej Karpathy's Software Is Changing (Again) argues for partial autonomy products with a human in the loop and an autonomy slider, and insists on the decade of agents rather than the year of them. That is the same argument from the product side that Lance makes from the implementation side, and the two videos agree on the mechanism: the slider Karpathy draws is the spectrum in Figure 1, and the thing that lets you move it is the interrupt and checkpoint machinery Lance names in his last ninety seconds. Ilya Sutskever on a decade of sequence to sequence learning is the research side of the same question, since whether Lance's forecast lands depends on whether the capability curve he is betting on keeps bending.

If you want one line to carry out of Part 4: reach for a workflow first, and build an agent only when the path genuinely cannot be written down.

Full transcript
[00:00:00] hi this is Lance from Lang chain and thropic recently put up this blog post called building effective agents where they Define what an agent is and also explain what a workflow is so I'm going to build every workflow and agent that they talk about in this blog post from scratch and explain what they all are how to build them and why each one can be very useful so here's a simple way to think about their definition of a workflow versus an agent think about a workflow as some kind of scaffolding of predefined code paths around LM calls [00:00:31] now this has been popular for years and it makes a lot of sense in many cases to take llm calls and embed them in some fixed set of code paths now sometimes you can actually have the LM decide what po what paths to take in a workflow and that's kind of this middle category that I drew here and we'll talk about all these in detail but that's kind of the intuition workflows are kind of like some scaffolding of predefined code paths and you can embed LM calls within it now agents remove the scaffolding so [00:01:02] you're basically let an llm direct its own actions now in this particular case actions are typically tool calls and an llm will receive the feedback from those tool calls and decide what to do next I do want to call out that you can certainly have llms that perform tool calls within workflows but the difference is that workflows have some scaffolding around them whereas an agent is unbounded so an agent just directly re receives the output of environmental feedback from [00:01:32] Tool calls decides what to do next without the kind of reasoning scaffolding around it so that's really the differentiation now just before I get started here it's important to call out why Frameworks okay so implementing these patterns doesn't require framework right sometimes this can be done in a few lines of code I completely understand that so you might say okay why are you talking about this why are you trying to promote Frameworks here's kind of my take on it Lang graph really means aims to minimize the overhead of implemen these patterns okay it doesn't abstract prompts it doesn't abstract [00:02:05] architecture what's really happening with langra is it's supporting infrastructure underneath any workflow or agent that's the key point and really it's three things one is persistence this gives you memory this gives you human the loops the ability to pause for example while an agent is processing and approve a tool call pause and riew what the agent is doing this is extremely useful in many cases streaming so of course LM apis stream tokens we all know that but when you build these worklow [00:02:36] rents sometimes you want to stream for example what certain steps output you might want to stream the outputs of tool calls you often want more flexibility over streaming than just what the llm calls produce and so Lang graph give you a lot of control over what you're outputting from your workflow or agent this is very useful when actually building these and putting them in production and the third is deployment so testing debugging deploying that is a big benefit of using Lang graph it is extremely easy to go from any workflow [00:03:06] or h&u Implement to a deployment lra doesn't abstract prompts and it does not abstract architecture it's really giving you low-level infrastructure that sits underneath any of these workflows or agents so now we've laid the foundations let's talk through the various patterns laid out in the blog post first starting with the augmented L LMS can be augmented with many different things one is for example M we talked a little bit about that that's one of the things that a framework such as landcraft can provide also LMS can interact with tools [00:03:36] and this is the foundation for building many workflows and agents so let's build this from scratch here's my notebook I've just pip installed a few packages set my in API key create my llm now let's show the augmentation for structured output I can take a schema in this case a pantic model I can bind it to my llm in this particular case I'm using Lang chain and I'm going to use the wi structured output method Lang chain to do this but again Lang graph does not require you to use Lang chain you can use the raw model apis it's completely fine I run that and the [00:04:08] output adheres to my schema which I passed here so that's great now let's do the same with tool calling I'm going to take a tool I'm going basically Define this function multiply as a tool that the LM has access to I can use the bind tools method in line chain to bind it to my llm now I have an llm with tools I invoke it with an input that's likely to elicit the tool call and I get the output great you can see a tool call is [00:04:38] produced it takes this input and creates the arguments necessary to actually run this function that's it remember when llms are creating tool calls they're really just giving you payloads to actually run that tool you could then run the tool pass the output of the tool back to llm and really doing that Loop gives you an agent like we saw before we've talked about augmented llm as a building block now let's talk about some of the workflow patterns covered in the blog post starting with prompt chaining here's the intuition each LM call [00:05:08] process the output of the previous one when do you want to do this when you ATT tested you can decompose into a few different llm calls so in this little example they show there's an input call one you can have some gating on the output of call one output goes to call two that output goes to call three so let's basically build this from scratch and let's do an example we build a chain that takes a topic from a user the LM makes a joke we check to make sure the joke has a punchline and we improve it [00:05:38] twice with two subsequent llm call now when I'm working L graph all I need to do is basically Define a container for everything that I want to modify over my workflow so in this particular case I'm just going to create a dict and it's going to contain topic that's we'll get from the user the joke will be the output of that first call improved joke will the output of the second call final joke would the output of the final call now when laying out workflows I actually like to draw them out anyway so this is kind of cool that the blog post has all these laid out and when you draw these [00:06:09] out think about for example each of these calls or steps is just a different function so for this particular workflow I'm going to create a function generate joke improve joke Polish joke and those are all going to be just simple llm calls now what's interesting is with Lang graph this container or state that I Define is pass into every one of these steps and I can extract whatever I want from it in this case I go State topic to get the topic that written to state by [00:06:39] the user and I can write things back to State just by basically returning this joke key the output of my LM call improve joke I populate improv joke final joke I populate final joke so what's happening is at each see these steps I'm making LM calls and I'm populating my state or this container in the workflow with the output of each llm [00:07:10] call that's really all it's going on now notice in this workflow there's a gate here so in this case I'm just going to create some gating logic so what's interesting is I can take in the state and I can just check hey does a joke have a question mark or an exclamation point that's just some arbitrary criteria and what I can do is I can pass anything I want here just two strings pass or fail now this can be used as conditional Edge or gate in Lang graph so now I have a container that has [00:07:41] everything I want to modify in my workflow I've defined each step of my workflow as an independent function I've defined a gate that's going to serve as that check on the output of the joke and now I lay this out in Lang graph as a simple workflow so all I need to do here is just take in my state initi ize the workflow add those three steps to it we added generate joke improve joke Polish joke and this is where I Define the connectivity of my graph or workflow so [00:08:11] I'm going to start I'm going to go to generate joke first and you can see then I want to check the output of my first call to see whether or not it has a punch line so I add a conditional Edge that connects generate joke and based upon the output of this function if it's pass I go to improve joke that's my other node if it's fail I just end then I create the edges from improve to Polish Polish to end compile that as a [00:08:42] workflow and there we go so that's the exact same workflow that they drew out in the blog post applied to joke generation now once I have this workflow I can very simply just run chain. invoke that's a very simple way to invoke any chain workflow or whatever whatever you have in Lang graph I pass in a topic from the user and we're can see the llms create an initial joke they improved it and Claude creates final joke so this is a very simple example of a chain great [00:09:12] so we've covered the augment to LM we've covered basic prom chaining now let's get into some more interesting and complex workflows like parallelization now let's talk about parallelization do this when you for example have multiple perspectives that you want for single task so I've this quite a bit with things like multi-query rag if I have a question I want to Fan it out into like three or four different sub questions or when independent task can be performed with different prompts or different llms so lots of cases where you want to just paralyze tasks and [00:09:42] workflows so let's just build this let's take a topic create a joke story and poem all in parallel so in the same way I did before I create my state again this is just a container for everything I'm going to modify in my workflow in this case I'll have a topic from a user and I'll have the joke the story the poem and all just combine them at the end now this workflow is going to have three LM calls they'll run in parallel and then one aggregation that'll pull them all together so all I'm going to do is I'm going to create a function for each of those different steps in my [00:10:13] workflow LM call one is going to write a joke about the topic a story a poem and I'll just aggregate them all into a string super simple now just like I showed before we can build this in L graph pass in that state the container add the nodes and just connect them so in this case I go from start to LM call 1 2 3 and they all connect to the aggregator on the end and I can show it there you go so you have a nice visualization of your workflow and I can run it there we go so we have a story [00:10:43] now a joke and a poem all created in parallel so we've covered the augmented elements our building block we talked about prom chaining we talked about parallelization now let's get into routing I've used routing a lot and it's extremely useful in a bunch of use cases when you want to decide for example if you want to only go to one of those steps which one to go to you can use LM to make that decision so a specific example this is I've used routing a lot in working with retrieval in taking a question and routing it to different retrieval systems so let's do [00:11:14] an example where we take an input and we route it to either a joke a story or a poem generation based on what the user asks for now I'm going to show a very nice trick here for routing I often like to do this I basically take an llm and I give it a structured output so that guarantees the LM is going to produce in this particular case the strings poem story or joke as a structured object and I'm going to give it in this particular case a pedantic model just like before I'm going to [00:11:44] Define my state again this is a container we'll take an input from the user we'll get the decision from our router and we'll get an output we'll save that as the final output now in this case for my nodes just like before four I'm going to have three different LM calls that will either write a story a joke or a poem and I'm going to have my router which is going to take the user input and basically decide whether or not to Route it to story joke or poem based on the content of the input and [00:12:14] it's just going to return that as a structured object and I can extract from the data model the step decision write that to State as decision so remember if you look at the model here it's just a pantic model which is going to Output a single key step which is either going to be poem story or jokes I extract that from the pedantic model that is returned by the router and I write that to state so then I have my decision in state and now this is just going to be a [00:12:44] conditional Edge that'll look at the decision and determine what node to go to super simple so basically if the decision is story go to node one joke 2 poem 3 that's it now you'll see in this toy example each of these steps are doing the same thing but in real world examples the router will send to different steps that will have different logic different LM calls this is just a toy example showing you how to hook up that logic between for example a router [00:13:16] step a structured output and a decision about where to go next cool so this is showing you visualization of that you'll see what's kind of nice in Lang graph when we visualize this this dotted line means a condition Edge so it's going to go to only one of those three paths whenever this runs and of course that's going to be based upon the decision of the router now we can look at these nodes and just create a print statement that'll just tell us which node was visited to confirm we can run this cool so we know we go to the node that [00:13:47] creates a joke and we get the joke output we've covered the augment elementer building block we talked about promp chaining we talked about parallelization we talked about routing and I'll talk about another workflow or rator worker now this is a really interesting one and I've used this quite a bit as well so this is a case where you want an llm to break down a task into a set of subtasks delegate each subtask to an independent worker and then synthesize the results so it's kind of like parallelization except the key difference is this worker assignment you don't know ahead of time so you're [00:14:18] having an llm reason about something and then create a bunch of workers based upon its reasoning so again in this case the LM is kind of gating or creating the control flow just like in the case of routing so here's an example that I've used quite a bit report writing maybe you've played with deep research and llm reasons about the plan for the report and dynamically generates a bunch of report sections and then goes and does research on all them classic example of an orchestrator worker type workflow so let's actually do that right now so for [00:14:49] this the trick is I'm also going to use structured outputs so I'm going to create a data model for a report section it's going to have a name and a description and I'm going to have a list of sections here so what I'm going to do is I'm going to bind that section list to my llm and that's going to be my planner so what's cool here is the planner is going to take an input reflect on it and produce a list of sections based upon Its Reflection so this is dynamic I don't know how many sections will create a priori that's why this is a very good orchestrator worker use case now in line graph look like [00:15:20] before I create my state now in this case what I'm going to do is I'm going to create a state for my orchestrator graph which is going to have like a topic from a user a list of sections a list of completed sections which the workers will all write to and a final report okay so that's State one now this is where things get interesting with orchestrated worker workflows in langra the way we often like to do it is for the workers give them their own state so why is this because each of those [00:15:50] workers you want to handle independent inputs and they're kind of all self-contained objects think about as their own little buckets in which work's being done and different Works being done in each one but they're all writing out to the same output and this is why I include this completed sections key in the worker State and in the graph state so what's interesting in line graph is when you have overlapping keys and you write to for example this completed [00:16:20] sections key in each worker the outer state will also have that update reflected so what's going to happen is all the worker is going to write to this completed sections key in parallel and we structure this key with an annotation that allows for the addition of new elements so that's really all we need to do so here I'm going to create a work an orchestrator and this is basically going to be my planner I'm going to invoke it with the input topic from the user and [00:16:51] I'm going to tell it to create a plan then I'm going to have this LM call which is my worker it takes in the worker state and it basically says write a report section and that's all I need and I just pass it the section name and section description and this completed sections is a state key that all the workers can write to in parallel that's the key point and all those sections are going to be accumulated in that completed sections key then I'm going to have a synthesizer that's going to read [00:17:21] out the completed sections and just write them all out as a string that's it write that to the final report key in my state now this is the only thing that's new and a little bit special in orchestrator worker style workflows because this is so common we have a special API called send and Lang graph that allows you to basically spawn these workers dynamically and what's happening is remember my planner wrote the sections of the port to the state and I can iterate through those sections [00:17:52] and I can basically assign each section to an independent worker just like this for each section in section s send to llm call that's my worker and basically initialize the state as section when I do that so then that LM call you can see receives worker State and it's receiving from State the section name and description and then it goes and writes that section it writes it output to [00:18:22] completed sections which my orchestrator has access to and then the synthesizer basically just grabs completed SE and combines them that's it you're done that's all you need to do now I can build this workflow and there we go so this dotted line just shows you that you use the send API to spawn a whole bunch of llm call workers you don't know how many a HEB ton that's why it doesn't draw out each one specifically it's going to be dynamically determined based upon the orchestrator plan that's the key characteristic of these orchestrator worker workflows you don't know a priori [00:18:53] how many workers you need the llm will determine that on the fly now let's Show an example this I want a report on L scaling logs so what's kind of cool is this runs fairly quickly you can see the planner generate all these report sections great and I can look at the final report and it concatenates them all together and you can see I get this Rich introduction and then I get you know fundamentals scaling relationships following the plan now in reality when I do this I have much more detailed [00:19:23] prompts I showed here this is really showing you the workflow rather than the particulars of How I build a high quality report writer in fact I've have separate videos on report writing you could check out but this is just showing you how to set up an orchestrator worker style of workflow if we just talk about the orchestrator worker workflow let's talk about something that's a bit related the evaluator Optimizer workflow in both these cases you're going to have llms directing the control flow through predefined code paths we have one llm [00:19:54] gener response and another kind of grade it and give feedback in a loop now I've used this quite a bit for example grading responses from a rag system for hallucinations or for factual accuracy I've used this kind of like kind of evaluator Gates and if for example there's hallucination in the output that isn't grounded by the documents I send it back and have a regenerate response and here I'm going to use structured outputs again you see there's kind of a trend here structured outputs is like kind of all you need and quotes it's extremely convenient because you can [00:20:24] really build all these workflows just of structured outputs you don't have to use tool calling for example you can use routers to basically conditionally determine where to go next you can have then nodes that for example just call tools depending on the result of the router itself so really just structured outputs is a nice way you can build a lot of complex workflows in this case I'm going to create a structured output that's basically my my grader model that's going to be grade and feedback decide if the joke is funny or not in my case and if it is not funny give some [00:20:56] feedback okay so that's going to be my evaluator and again I'm going to find the graph state in this case I'm going to take a joke based upon a topic from a user I'll generate the joke and then I'll grade it is give it feedback determine if it's funny or not based on the feedback I'll go back regenerate a new joke that's it nice and simple so this is going to be my generator it's going to generate a joke now you'll see some do something kind of interesting here I'm going to check if there's feedback in the state okay now there might be because I've basically looped [00:21:27] back to this node if I determine that the joke is not good so there may be feedback in the state if there's feedback included in my prompt otherwise I just say write a joke about the topic so nice and easy and then I have an evaluator that basically takes in the joke from State and grades it again this evaluator has structured output so it'll basically produce a grade and some feedback and then I have additional Edge that'll look at State funny or not so again that's like kind of my grade and if it's funny route to accepted if it's [00:21:59] not funny route to rejected and feedback so again this is the initial ledge I use to determine where to go next just like we saw it's routing build My Graph and you can see has everything we want here here's the generator here is the evaluator based on the evaluator we either go back and again you can see we set that conditional ledge up right here so that route joke conditional ledge we just talked about it if funny or not if it's funny we go to accepted so if accepted we go to end if it is rejected [00:22:29] in feedback we go back to LM call generat so you can see when you st the initial Edge this is how you can basically route from the output of your logic your Edge logic to the next node to go to that's it it's extremely simple and you can see you get a nice visualization of it so let's give this a shot run it with an input of cats so we get the out of the joke and we actually can look at the state to see what the feedback was so State feedback in this case it seems to like it and let's check [00:23:02] funny or not and it determines it's funny okay so greater like the joke we went ahead and returned it then to the user we ended and so we can see that the greater was initiated and decided to like the joke so it passes and we finish so we talked about a number of different workflows that use llms within some kind of reasoning scaffolding and in the case of orchestrator worker evaluator Optimizer routing you actually do let the L M make decisions to Route the control flow through that scaffolding [00:23:32] now let's remove the scaffolding and let's talk about agents with agents You're simply allowing an llm to form actions in form of tool calls and directly receive the output or feedback from those actions and so in the workflow case we talked about there are always kind of these predefined code paths that we had an llm kind of follow and kind of route through in the case of an agent we've removed those now when do you actually need an agent this is kind of the big question you see [00:24:03] agents being used in cases where you really have open ended problems that you cannot easily capture in a workflow for example you want llm to utilize different Tools in a pattern that you just cannot predict out priority so it's not easy to lay it in the workflow so it's kind of open-ended task we some some really interesting examples of challenges like s bench so a benchmark for software engineering in which anthropic actually used a agents architecture just shown like this and [00:24:33] achieves very strong performance so we know for certain open-ended tasks agents are appropriate I do want to caveat in the event that LMS get extremely proficient at tool calling it's also possible that a lot of the scaffolding that we talked about with various workflows is unnecessary today if you know roughly the sequence tools need to be initiated it's often better just capture it in a workflow in terms of reliability than just give it to an agent let the agent [00:25:03] hopefully call that correct sequence now let's just set up an LM with tools I'm going to give it multiply add and divide nice and simple now I'm going to Define three nodes so I'm going to call my LM and I'm going to basically allow the llm to call a tool okay so I'm using my LM with tools and in this particular case I'm saying you're a helpful assistant task with performing arithmetic so the output of that tool call is going to be saved to this messages key in our state so what's happening is my state has a [00:25:33] single key in this particular place messages which is going to accumulate what I pass as the user what the LM produces and so forth so what's going to happen is I have another node called tool node it's going to look at the state look at the last message determine if it's a tool call if it is it'll go ahead and just call that tool that's it nice and simple and it's going to return that to state so what's interesting here is I'm G to have a sequence of human [00:26:03] input model agent in this case the size to call a tool this tool node looks and sees oh the LM decided to call a tool it actually runs that tool call that then is written to State messages as a tool message this is that environmental feedback thing you hear about so you talk about agents they can perform actions that's the tool call done up here fine they also can receive feedback from the environment and act on it that is the output of this tool noes this tool Noe is basically the environmental [00:26:34] feedback saying here's the output of the tool call the llm then we'll get that decide what to do next that's it that's an agent now the only other thing I need is this conditional ledge that basically says was the last message uh a tool call if so I'm going to go ahead and route to the tool node and if not I will end so you know you can modify this in different ways but a lot of times people basically just say allow the agent continue making tool calls until it [00:27:04] decides it doesn't need one anymore and then you don't and then you're done so there we are I mean that's our agent Loop that you know and that's kind of why agents are elegant they're extremely simple in this kind of formulation all is it's basically an llm initially with a bunch of tools and it's like a tool node that will run the tool for you and return that environmental feedback to the llm and let the llm keep spinning until it decides you don't need a tool call anymore then you're done that's really it now again this tool call thing think about that is actions so this LM can [00:27:34] perform actions the actions are determined by the user so you can pass in any set of tools to this LM it decides to perform those actions now you need some system that actually does those actions so that's like that's what we create with this tool node and this just Loops until the LM says I don't need a tool call anymore and then that conditional Edge says okay I can just end that's really it so let's see an example of that I'm going to basic tell this agent add three and four then take the output and multiply by four okay now [00:28:05] look this is extremely simple and LM can just do this I totally get it this is more showing you the principle setting up an agent with these tools for Edition and multiplication and testing whether or not it can correctly perform these tool calls in sequence and looking at the flow of messages so here we go this is pretty cool now again this is like a toy example Le but that's showing you the flow and that's what matters you have an input from the human my instructions the llm looks at that and says okay I need [00:28:35] to make a tool call makes a tool call the tool node executes the tool call returns environmental feedback to my llm as the result of the tool call which is seven my agent thinks about it makes another tool call responds with 28 agent thinks about it final result is 28 here's how we got there no more tool calls needed done for all the hyp agents that's all it is extremely simple is tool calling in a loop now again why don't we just do this because look a lot of problems can be solved with workflows [00:29:05] which are bit simpler agents haven't been particularly reliable to date particularly with large numbers of tools or complex trajectories of tool calls and so blot people actually in production prefer workflows I've done a lot more workflows than agents to be honest over the last year or two but I completely acknowledge in the event that you have ver capacity reasoning models that can perform low latency high quality tool calling and we have more kind of confidence that they can perform reliably in production I think you will see the movement to [00:29:35] this classic style of very simple tool calling agent in production but the game in the field today is a little bit more like people are saying well I want to put workflows into production because they're a little bit more trustworthy again I have some scaffolding around the core kind of L calls now I do also want to show this is looking at the documentation that we do have AE pre-built method called create react agent that basically wraps what we just built from scratch right here for convenience now this depends if you want [00:30:07] to build it from scratch yourself just as we did that is completely fine if you want to use the preo method that's fine as well but it's available to you and I will be sharing the link to this tutorial page which has everything we just went through so all the different uh workflows we talk through are all in this tutorial which I will be sharing but just from this little tutorial you've seen here's how you can lay out all these different workflows and an agent all in Lang graph if I go back to the Y langra story when [00:30:38] you get with langra is when you compile those workflows or agent in langra you're getting a persistence layer for free which gives you short and long-term memory and that also gives you the ability to stop perform interruptions review and continue IE human in the loop it gives you a whole bunch of streaming capacities you can stream independent values from your State at any point in time you can stream of course tokens out of LM calls and you also get deployment so we have a very nice and easy onramp for testing debugging and deploying any [00:31:08] of these so you could take any of those workflows R we just built and deploy them in like five minutes so very very quickly that's really what you're getting with the framework and you can also see with L graph it's pretty easy to lay them out and actually working on making it even easier so we're trying to reduce the overhead so you can lay it out almost as if you're writing python you're not even thinking about the framework but you're getting these benefits kind of for free when you do use a framework that's kind of the big idea here so hopefully this was a helpful overview to present uh how to [00:31:39] lay these various workflows or agents out using Lang graph and what benefits you may get from L graph as a consideration so thanks very much feel free to leave any comments below