local-ai·lab
Lesson 8

A Stateful Agent with LangGraph

Turn the linear RAG pipeline into an agent that grades its own retrieval, rewrites its query and retries - then price the graph honestly against the same loop written as sixty-three lines of while. Three arms over Lesson 7's corpus: the chain fixes 4 of 8, the loop fixes 8 of 8 and correctly refuses the ninth, and the graph matches the loop exactly. Everything LangGraph costs buys checkpoints, interrupts and an inspectable topology - not better answers. Python and Node.js.

Download lesson PDF PDF · downloadable
Follow along in:
Overview

What you'll build

New to LangGraph? It is not LangChain 2.0, and it is not an agent framework in the "hand it tools and hope" sense. It is a small state-machine library with checkpoints: you declare a typed state, write nodes that return updates to it, and connect them with edges - including edges that point backwards. It ships Python and JavaScript SDKs, both used here, and it depends on langchain-core, so this lesson sits on top of Lesson 7's dependency rather than beside it.

Lesson 7 rebuilt the pipeline on a framework and asked what it cost. This lesson asks the harder version of that question. A corrective loop - retrieve, grade the evidence, rewrite the query, try again - fixes four of the questions here that a single retrieval gets wrong, and correctly refuses a fifth that nothing in the corpus answers. But you do not need LangGraph for a loop. You need while.

So this lesson runs three arms, not two: the linear chain, the same loop written as sixty-three lines of while, and the loop as a StateGraph. The while loop and the graph agree on every single question. That zero is the lesson. Pick a language above and press → to begin.

   A chain runs once. A graph decides whether to run again.

                        ┌───────────────────────────────┐
                        │  START     question, attempt=1 │
                        └───────────────┬───────────────┘
                                        ▼
                        ┌───────────────────────────────┐
              ┌────────▶│  retrieve     BM25 · top-k     │  retrievals +1
              │         └───────────────┬───────────────┘
              │                         ▼
              │         ┌───────────────────────────────┐
              │         │  grade      coverage or LLM    │  grades +1
              │         └──────┬─────────────────┬──────┘
              │      weak, and │                 │ grounded
              │   attempts left│                 │
              │                ▼                 ▼
              │  ┌──────────────────────┐   ┌──────────────────────────┐
              └──┤  rewrite   glossary  │   │  review     ⏸ interrupt()│
                 └──────────┬───────────┘   └───┬──────────┬───────────┘
                            │       approve     │          │ veto
                    edit    └───────────────────┤          │
                                                ▼          ▼
   weak, no attempts left            ┌──────────────┐  ┌──────────────┐
   ─────────────────────────────────▶│   abstain    │  │   generate   │
                                     └──────┬───────┘  └──────┬───────┘
                                            ▼                 ▼
                                           END               END

   Three things this picture has that a chain cannot draw:

     ◀──   an edge that points BACKWARDS            (rewrite → retrieve)
     ⏸     a node that STOPS and waits for a human  (review)
     ▼ ▼   two different ways to finish             (generate · abstain)


   Measured over Lesson 7's corpus, nine questions:

                              linear      loop (while)   graph (LangGraph)
     top source correct          4/8               8/8                 8/8
     correctly abstained         0/1               1/1                 1/1
     retrieval calls               9         15 (+67%)                  15

   The while loop and the graph agree on every question. That is not a
   disappointment - it is the finding. Everything LangGraph costs is paid
   for the three things the loop cannot do at all: keep state past the
   process, stop halfway and ask a human, and tell you its own shape.
Setup

What you need

This is the second lesson that installs something, and like Lesson 7 the dependency is part of the argument. ./run -l 8 puts langgraph into the course venv on first use. If you skip it, the demo still runs the linear and loop arms, prints the whole accuracy result, marks the graph column not installed, and exits 0 - which is a better degradation than Lesson 7 managed, because here the dependency is not what produces the result. Run everything from the repo root:

run
$ ./run -l 8                    # Python: the playground (default)
./run -l 8 demo               # Python: the three-arm comparison and the scorecard
./run -l 8 --lang node demo   # Node.js: same control flow, a different bill
Step 1

Nine questions, and why each one is here

Three are carried straight from Lesson 7 as controls. Four are phrased the way a real person asks - orange light when the docs say amber status ring, wipe when they say factory reset, RMA when they say warranty claim, 5GHz when they spell out five gigahertz. One is already correct on the first try and makes the loop spend a retrieval for nothing. One has no answer in the corpus at all.

Every question carries its own why, and the demo prints it, so the reason a question exists lives in the data rather than only in this prose.

data/questions.json
{
  "chunk_size": 700,
  "chunk_overlap": 120,
  "top_k": 3,
  "grade_threshold": 0.67,
  "max_attempts": 3,
  "questions": [
    { "id": "q1", "ask": "What is the factory reset procedure?",
      "expect": "manual.md",
      "why": "control - carried from Lesson 7, and it must not regress" },
    { "id": "q2", "ask": "Why does the status ring stay amber?",
      "expect": "faq.md",
      "why": "control - carried from Lesson 7" },
    { "id": "q3", "ask": "Which endpoint exports the logging buffer?",
      "expect": "api.md",
      "why": "control - carried from Lesson 7" },
    { "id": "q4", "ask": "The light is orange and it will not connect.",
      "expect": "faq.md",
      "why": "vocabulary - the reader says 'orange light', the docs say 'amber status ring'" },
    { "id": "q5", "ask": "How do I wipe the device and start over?",
      "expect": "manual.md",
      "why": "user-speak - the reader says 'wipe', the docs say 'factory reset'" },
    { "id": "q6", "ask": "What paperwork does an RMA need?",
      "expect": "warranty.md",
      "why": "acronym - the reader says 'RMA', the docs say 'warranty claim'" },
    { "id": "q7", "ask": "Is 5GHz supported?",
      "expect": "networking.md",
      "why": "notation - the reader types '5GHz', the docs spell out 'five gigahertz'" },
    { "id": "q8", "ask": "Can I mount it sideways?",
      "expect": "installation.md",
      "why": "FALSE ALARM - already correct on the first try; the grader loops anyway and buys nothing" },
    { "id": "q9", "ask": "What is the MTBF?",
      "expect": null,
      "why": "not in the corpus at all - the loop has to give up rather than spin" }
  ]
}
Corrective RAG only demonstrates itself on questions the first search actually fails. These fail for a boring, reproducible reason - the words are not in the documents - which is exactly what makes the result provable offline, through retrieval alone, with no model in the loop.
Step 2

The grader - one judgement the whole graph turns on

Term coverage: what fraction of the query's content words appear anywhere in the retrieved text? It is a real information-retrieval signal, it detects precisely the failure this lesson is about, and missing - the terms that did not land - is exactly what the rewriter needs next.

python/graders.py
class CoverageGrader:
    """Term coverage: what fraction of the query's content words does the evidence contain?

    This is a real information-retrieval signal, not a stand-in for one. It
    detects the exact failure this lesson is about - a reader asking in words the
    documents never use - because those words cannot appear in any retrieved
    chunk, whatever the retriever ranks first.

    It also explains itself. `missing` is the list of terms that did not land,
    which is precisely what the rewriter needs, so the two nodes are coupled
    through data rather than through a shared hard-coded table.
    """

    name = "coverage"

    def __init__(self, threshold: float) -> None:
        self.threshold = threshold

    def grade(self, question: str, query: str, docs: List[rag_core.Chunk]) -> Grade:
        terms = rag_core.content_terms(query)
        if not terms:
            # A query with no content words cannot be judged this way. Say so
            # rather than dividing by zero or silently passing.
            return Grade(verdict="weak", score=0.0, missing=[],
                         reason="no content terms to look for", grader=self.name)
        haystack = set(rag_core.tokens(rag_core.evidence_text(docs)))
        missing = [t for t in terms if t not in haystack]
        score = (len(terms) - len(missing)) / len(terms)
        grounded = score >= self.threshold
        return Grade(
            verdict="grounded" if grounded else "weak",
            score=round(score, 2),
            missing=missing,
            reason=(f"evidence covers {len(terms) - len(missing)}/{len(terms)} query terms"
                    + ("" if grounded else f"; missing {missing}")),
            grader=self.name,
        )
The grader is never shown the expected answer. It cannot be: expect lives in the data and is used only to score the run afterwards. A test asserts that on the function signature, because a grader that can read the label turns the whole measurement into theatre.
Step 3

The rewriter - and an admission

When the grader says the evidence is weak, the query has to change or the next attempt just repeats the last one. This expands the missing terms into the words the documents actually use, and drops the ones that landed nowhere.

python/rewriter.py
class GlossaryRewriter:
    """Expand the terms the evidence did not cover into the words the docs use.

    Only `missing` terms are expanded. Rewriting on a grade with nothing missing
    is a no-op, which keeps the rewriter honest: it reacts to the grader's finding
    instead of reaching for the table whenever it is called.
    """

    name = "glossary"

    def __init__(self, glossary: Optional[Dict[str, List[str]]] = None) -> None:
        self.glossary = glossary if glossary is not None else rag_core.load_glossary()

    def rewrite(self, question: str, query: str, missing: List[str]) -> str:
        if not missing:
            return query
        expanded: List[str] = []
        for term in missing:
            expanded.extend(self.glossary.get(term, []))
        if not expanded:
            # Nothing in the table matches. Returning the query unchanged is the
            # honest move: the loop will grade it weak again and, at the cap,
            # abstain - which is the correct outcome for a question the corpus
            # does not answer. Inventing a rewrite here would turn "we have no
            # document about this" into "we tried harder", which is worse.
            return query
        # Keep the terms that DID land, drop the ones that did not, add the
        # documents' words. Dropping the misses is the point: "orange" is not in
        # any chunk, so leaving it in only dilutes the BM25 score of the rest.
        kept = [t for t in rag_core.content_terms(query) if t not in missing]
        out: List[str] = []
        for term in kept + expanded:
            if term not in out:
                out.append(term)
        return " ".join(out)
It is a lookup table, and a lookup table is a fixture rather than a design. It fixes orange only because somebody wrote orange into a JSON file, and it will not fix the phrasing nobody anticipated. It exists so this page prints the same thing on every machine. That is the entire argument for --llm-rewrite, and the lesson would rather say so than let you discover it.
Step 4

Arm B - the objection, written out in full

Before any graph: here is the corrective loop as a plain while. It calls the same retrieve, the same grader and the same rewriter the graph will call. Note the early exit - if the rewriter hands back the query unchanged, there is nothing left to try, and a cycle that cannot change its own input is an infinite loop with extra steps.

python/loop_agent.py
def run(retriever, question: str, *, grader: Grader, rewriter: Rewriter,
        top_k: int, max_attempts: int,
        generate: Optional[Callable] = None) -> dict:
    """Retrieve, grade, and re-search with a better query until the evidence holds up."""
    query = question
    attempt = 1
    retrievals = grades = rewrites = 0
    trace: List[str] = []
    verdicts: List[str] = []
    docs: List[rag_core.Chunk] = []
    grade = None

    while True:
        docs = rag_core.retrieve(retriever, query, top_k)
        retrievals += 1
        trace.append(f"retrieve  attempt={attempt} query={query!r} -> {rag_core.sources(docs)}")

        grade = grader.grade(question, query, docs)
        grades += 1
        verdicts.append(grade["verdict"])
        trace.append(f"grade     {grade['verdict']} score={grade['score']} - {grade['reason']}")

        if grade["verdict"] == "grounded":
            break
        if attempt >= max_attempts:
            trace.append(f"abstain   {attempt} attempts used, evidence still weak")
            break

        new_query = rewriter.rewrite(question, query, grade["missing"])
        rewrites += 1
        if new_query == query:
            # The rewriter had nothing left to try. Retrying an identical query
            # would retrieve identical chunks and grade identically - a cycle that
            # cannot change its own input is an infinite loop with extra steps.
            # Stopping here reaches the same decision the cap would, one wasted
            # retrieval sooner.
            trace.append("abstain   rewrite changed nothing - no point asking again")
            break
        # A rewriter that quietly fell back to the glossary must not look like one
        # that produced this query itself.
        note = "  [fell back to the glossary]" if getattr(rewriter, "fell_back", False) else ""
        trace.append(
            f"rewrite   {query!r} -> {new_query!r}  (missing {grade['missing']}){note}")
        query = new_query
        attempt += 1

    grounded = grade["verdict"] == "grounded"
    answer = ""
    if grounded and generate is not None:
        answer = generate(question, docs)
    elif not grounded:
        answer = rag_core.ABSTAIN_TEXT

    return {
        "question": question, "query": query, "attempts": attempt,
        "retrievals": retrievals, "grades": grades, "rewrites": rewrites,
        "sources": rag_core.sources(docs) if grounded else [],
        "verdicts": verdicts, "grade": grade, "docs": docs,
        "status": "answered" if grounded else "abstained",
        "answer": answer, "trace": trace,
    }
Benchmarking LangGraph against a single-shot chain would prove that looping helps, which is not the same claim. A reader would be right to answer "so write a while loop". Lesson 7's authority came from refusing to benchmark against a straw man; dropping that standard one lesson later would be worse than never having had it.
Step 5

State, and what a reducer actually is

Look at the two kinds of field in one struct. trace, turns and the three counters are Annotated[..., operator.add], so a node returning {"retrievals": 1} adds one to the running total. Everything else is last-write-wins, so a node returning {"query": ...} replaces it.

python/graph_agent.py
class AgentState(TypedDict, total=False):
    """Everything that flows between nodes.

    Note the two kinds of field, deliberately side by side:

      `trace`, `turns`, and the three counters are Annotated with operator.add,
      so a node returning {"retrievals": 1} ADDS ONE to the running total.

      everything else is last-write-wins, so a node returning {"query": "..."}
      REPLACES the query.

    That difference is what a reducer is, and it is easier to see in one struct
    than in any amount of prose.
    """

    question: str          # the human's words. Never mutated.
    query: str             # the CURRENT search query. `rewrite` replaces this.
    attempt: int
    docs: List[rag_core.Chunk]
    sources: List[str]
    grade: Grade

    trace: Annotated[List[str], operator.add]
    turns: Annotated[List[dict], operator.add]
    retrievals: Annotated[int, operator.add]
    grades: Annotated[int, operator.add]
    rewrites: Annotated[int, operator.add]

    decision: str          # "approve" | "veto" | "edit:<new query>"
    answer: str
    status: str            # "answered" | "abstained" | "vetoed"
That contrast is the clearest available definition of a reducer, and it is far easier to see side by side in one TypedDict than in any amount of explanation. It is also why the scorecard's counters are measured rather than asserted: the graph counts itself.
Step 6

Conditional edges, and the edge that points backwards

add_conditional_edges takes a function that returns a string key, and a map from those keys to nodes. rewrite -> retrieve is the cycle. grade has three or four ways out depending on whether a human is in the loop.

python/graph_agent.py
    def route_after_grade(state: AgentState) -> str:
        if state["grade"]["verdict"] == "grounded":
            return "review" if human_review else "generate"
        if state["attempt"] >= max_attempts:
            # The cap is a decision, not a crash. It produces a good answer -
            # "not in your documents" - rather than an exception.
            return "abstain"
        return "rewrite"

    def route_after_review(state: AgentState) -> str:
        decision = (state.get("decision") or "approve")
        if decision == "veto":
            return "veto"
        if decision.startswith("edit:"):
            return "retrieve"
        return "generate"

    builder = StateGraph(AgentState)
    builder.add_node("retrieve", retrieve_node)
    builder.add_node("grade", grade_node)
    builder.add_node("rewrite", rewrite_node)
    builder.add_node("generate", generate_node)
    builder.add_node("abstain", abstain_node)
    if human_review:
        builder.add_node("review", review_node)
        builder.add_node("veto", veto_node)

    builder.add_edge(START, "retrieve")
    builder.add_edge("retrieve", "grade")
    targets = {"rewrite": "rewrite", "generate": "generate", "abstain": "abstain"}
    if human_review:
        targets["review"] = "review"
    builder.add_conditional_edges("grade", route_after_grade, targets)
    builder.add_conditional_edges("rewrite", _rewrite_router,
                                  {"retrieve": "retrieve", "abstain": "abstain"})
    if human_review:
        builder.add_conditional_edges("review", route_after_review,
                                      {"generate": "generate", "retrieve": "retrieve",
                                       "veto": "veto"})
        builder.add_edge("veto", END)
    builder.add_edge("generate", END)
    builder.add_edge("abstain", END)

    graph = builder.compile(checkpointer=checkpointer)
    graph._retrieve_node = retrieve_node  # so run() can attach the retriever
    return graph
Because the routing function returns a key rather than a node, the topology stays separable from the logic - which is what lets get_graph().edges print the back-edge as a value. A while loop cannot answer "what are your edges?" at runtime. Its control flow is if and while; this one is data.
Step 7

interrupt() - stopping to ask a human

The graph stops here and hands back a review packet: the question, the query it ended up searching, how many attempts it took, the citations, and the evidence itself. You resume with Command(resume="approve"), "veto", or "edit:<a better query>" - and that last one re-enters the cycle at retrieve.

python/graph_agent.py
    def review_node(state: AgentState) -> dict:
        """Stop, and hand a human everything they need to judge the answer.

        The payload IS the review packet. That is why this uses the dynamic
        `interrupt()` rather than `interrupt_before=["generate"]`: the static form
        pauses without telling the reviewer anything, and a human approving a
        grounded answer needs to see the evidence it is grounded in.
        """
        decision = interrupt({
            "question": state["question"],
            "query_used": state["query"],
            "attempts": state["attempt"],
            "citations": state["sources"],
            "evidence": [d["text"][:200] for d in state["docs"]],
            "ask": "approve | veto | edit:<new query>",
        })
        out = {"decision": decision, "trace": [f"review    human said {decision!r}"]}
        if isinstance(decision, str) and decision.startswith("edit:"):
            # The reviewer did not just approve or refuse - they handed back a
            # better query. That re-enters the cycle at `retrieve`, which is a
            # second way into the loop and the reason `review` earns a node
            # rather than being a flag on `generate`.
            edited = decision[len("edit:"):].strip()
            if edited:
                out["query"] = edited
                out["attempt"] = state["attempt"] + 1
                out["trace"] = [f"review    human rewrote the query -> {edited!r}"]
        return out
This uses the dynamic interrupt() rather than the older interrupt_before=["generate"], because the payload IS the point. A static pause tells the reviewer nothing, and a human approving a grounded answer needs to see what it is grounded in. The edit branch is also why review earns a node instead of being a flag on generate: it is a second way into the loop.
Command

Run the comparison

demo · python
$ python -m pip install -q -r requirements.txt
demo · python
$ ./run -l 8 demo
demo · node
$ npm --prefix node install --silent --no-audit --no-fund
demo · node
$ ./run -l 8 --lang node demo
Run it

Read the output

The headline: 4/8 becomes 8/8, plus the abstention the chain never makes, for 67% more retrieval calls. The four the loop fixes are exactly the four where the reader's words are not the documents' words. q8 was already right and looped anyway - a retrieval that bought nothing, left in because pretending the loop is free would be a lie. q9 has no answer in the corpus, and the loop refuses rather than spinning.

Then the part that matters. The while loop and the graph agree on all nine questions - exactly. LangGraph did not make the agent smarter. Everything it cost bought the three rows the loop leaves blank: state that survives the process, a run you can pause and resume, and a topology you can print. The line counts and package counts are read off disk at run time, so they cannot drift from the code.

Lesson 8 · A stateful agent with LangGraph
corpus: lessons/07-langchain-rag/data/corpus  (7 documents - Lesson 7's, by path)
settings: top_k=3 threshold=0.67 max_attempts=3 grader=coverage rewriter=glossary

Nine questions. Three arms. The same corpus and the same settings.

       linear  loop    graph   att  question
  ------------------------------------------
  q1   PASS    PASS    PASS    1    What is the factory reset procedure?
  q2   PASS    PASS    PASS    1    Why does the status ring stay amber?
  q3   PASS    PASS    PASS    1    Which endpoint exports the logging buffer?
  q4   FAIL    PASS    PASS    2    The light is orange and it will not connect.
  q5   FAIL    PASS    PASS    2    How do I wipe the device and start over?
  q6   FAIL    PASS    PASS    2    What paperwork does an RMA need?
  q7   FAIL    PASS    PASS    2    Is 5GHz supported?
  q8   PASS    PASS    PASS    2    Can I mount it sideways?
  q9   FAIL    PASS    PASS    2    What is the MTBF?

  why each question is here
    q1  control - carried from Lesson 7, and it must not regress
    q2  control - carried from Lesson 7
    q3  control - carried from Lesson 7
    q4  vocabulary - the reader says 'orange light', the docs say 'amber status ring'
    q5  user-speak - the reader says 'wipe', the docs say 'factory reset'
    q6  acronym - the reader says 'RMA', the docs say 'warranty claim'
    q7  notation - the reader types '5GHz', the docs spell out 'five gigahertz'
    q8  FALSE ALARM - already correct on the first try; the grader loops anyway and buys nothing
    q9  not in the corpus at all - the loop has to give up rather than spin

The scorecard
                                   linear (L7)      loop (while) graph (LangGraph)
  top source correct                       4/8               8/8               8/8
  correctly abstained                      0/1               1/1               1/1
  --------------------------------------------------------------------------------
  retrieval calls                            9        15  (+67%)                15
  grade calls                                0                15                15
  rewrite calls                              0                 7                 7
  --------------------------------------------------------------------------------
  state survives a process                   -                 -      checkpointer
  pause / resume mid-run                     -                 -       interrupt()
  topology you can query                     -                 -       get_graph()

  The loop fixed 4 questions the chain got wrong (q4, q5, q6, q7), and
  refused the 1 it could not answer at all - for 67% more retrieval calls.
  q8 was already right on the first try and looped anyway: a retrieval that bought nothing.

  The while loop and the graph agree on all 9 questions: exactly.
  LangGraph did not make the agent smarter. It made it resumable,
  pausable and inspectable - the three rows above that only it can fill.

The topology, as data
  nodes  ['abstain', 'generate', 'grade', 'retrieve', 'review', 'rewrite', 'veto']
  edges
    __start__  -> retrieve
    abstain    -> __end__
    generate   -> __end__
    grade      -> abstain
    grade      -> generate
    grade      -> review
    grade      -> rewrite
    retrieve   -> grade
    review     -> generate
    review     -> retrieve
    review     -> veto
    rewrite    -> abstain
    rewrite    -> retrieve   <- the cycle
    veto       -> __end__

  A while loop cannot answer the question 'what are your edges?' at runtime.
  Its control flow is `if` and `while`; this one is a value you can print.

Human in the loop
  question: 'How do I wipe the device and start over?'
  This question matters: the manual says a factory reset destroys the logging
  buffer permanently, and the warranty page says a claim without that export
  cannot be assessed. Answering helpfully and immediately costs the reader
  their evidence. So the graph stops and asks.

  invoke #1 ->  the graph STOPPED. It did not answer.
                status              paused
                get_state().next    ('review',)
                answer in state?    False
                review packet       citations=['manual.md:1', 'troubleshooting.md:1']
                                    attempts=2  ask='approve | veto | edit:<new query>'
                retrievals so far   2

  invoke #2 ->  Command(resume='approve') on the same thread_id
                retrievals now      2      <- unchanged. It resumed; it did not restart.
                status              answered

  invoke #2' -> Command(resume='veto') on a fresh thread
                status              vetoed - generate never ran

  invoke #2'' -> Command(resume='edit:...') - the reviewer hands back a query
                retrievals now      3      <- it re-entered the cycle at retrieve
                a second way into the loop, which is why review is a node

Memory - two turns on one thread
  thread 'aurora'  turn 1  What is the factory reset procedure?
                   cited ['manual.md:1', 'warranty.md:1']
  thread 'aurora'  turn 2  Which endpoint exports the logging buffer?
                   cited ['api.md:1', 'manual.md:1']
  thread 'somebody-else'  turn 1  <- a different thread_id remembers nothing
  checkpoints written on 'aurora': 10

  MemorySaver keeps this in the process. For state that outlives the process:
    pip install langgraph-checkpoint-sqlite
    ./run -l 8 chat --sqlite /tmp/aurora.db --thread aurora "..."
  It is a separate package on purpose - see the bill below.

What it cost
                                        linear      loop (while) graph (LangGraph)
  code you maintain                   68 lines         298 lines         470 lines
  requirements lines                         0                 0                 1
  packages installed                        48                48                71
  install size                          ~34 MB            ~34 MB            ~43 MB

  The corrective loop cost 230 lines and ZERO packages, and it delivered
  the entire accuracy gain in the scorecard above. Most of those lines are the
  grader and the rewriter, which the graph needs too. The loop ITSELF is
  63 lines of `while`.

  Swapping those 63 lines for LangGraph cost 172 lines of graph
  wiring and 23 packages - for answers that are identical on all nine
  questions. What it bought is the three rows the loop could not fill:
  checkpoints, interrupts, and a topology you can query.
  If you already did Lesson 7 it is only 5 more packages (67 -> 72),
  because LangGraph depends on langchain-core and you are already paying for it.
  Which of those two numbers applies to you is the whole question.

grader: coverage (deterministic). Run with --llm-grade for the grader you
should actually use; it will not print the same thing twice, and that is the point.
Try it

The question where the words are wrong

Nothing in the corpus says orange, and nothing says light. The documents say amber status ring. The first search returns warranty.md and networking.md - plausible, and wrong. Watch the grader score it 0.00, name what is missing, and watch the second attempt land on faq.md.

The light is orange and it will not connect.
Try it

The question with no answer

The rewriter does its job correctly - it expands the acronym to mean time between failures - and the search still fails, because the corpus genuinely does not cover it. That is the distinction worth having: a bad query and a missing document are different problems, and only one of them is worth retrying. The agent abstains.

What is the MTBF?
Try it

Watch one question, attempt by attempt

Every retrieve, every grade, every rewrite, in order. Add --recursion-limit 4 to a question that loops and watch LangGraph's own cap fire instead of yours - they are not the same thing, and the lesson explains why that distinction matters more than it looks.

trace · python
$ ./run -l 8 trace "What paperwork does an RMA need?"
Try it

Two questions, one thread

Several turns on the same thread_id, so the checkpointer has something to remember. With no questions it asks two of the demo's; add --sqlite FILE and run it twice to watch the state survive the process.

chat · python
$ ./run -l 8 chat --thread aurora "What is the factory reset procedure?"
Try it

Stop the agent and decide yourself

This pauses mid-graph on real stdin and shows you the review packet before anything is generated. The question is not chosen at random: the manual says a factory reset destroys the logging buffer permanently, and the warranty page says a claim without that export cannot be assessed. Answering promptly and helpfully costs the reader their evidence. Approve it, veto it, or hand back a better query.

review · python
$ ./run -l 8 review "How do I wipe the device and start over?"
Try it

The grader you should actually be using

Run this one. It calls the LLM grader several times against byte-identical evidence and prints the spread of verdicts it gives back, next to the deterministic grader answering the same thing every time.

The deterministic grader is the default here for exactly one reason: demo is committed to a file and diffed by a test, so it has to print the same thing twice. That is a property this lesson needs and your system almost certainly does not. On your own documents, use --llm-grade - it is the better grader, it generalises to phrasings no glossary anticipated, and the fact that it will not repeat itself is something to measure rather than something to avoid.

spread · python
$ ./run -l 8 spread "Is 5GHz supported?" --runs 7
Try it

A real answer, through the graph

The comparison is deliberately offline. This runs the whole graph with a model at the end - retrieval, grading, rewriting, generation - through the course's configured provider. Add --llm-grade to put the model in the grading node too, which is how you would actually run this.

ask · python
$ ./run -l 8 ask "Why is the light orange?" --llm-grade
Try it

Print the graph's own shape

The topology as data: every node, every edge, and the back-edge marked. This is the thing a while loop cannot answer about itself at runtime, and it costs one call to get_graph().

graph · python
$ ./run -l 8 graph
Try it

Re-derive the dependency numbers here

Do not trust the bill in this lesson - recompute it. This walks LangGraph's actual dependency closure on your machine and counts the code off disk.

Look for langsmith in that list. LangChain's hosted tracing client arrives as a transitive dependency whether you asked for it or not. It is inert until you set LANGSMITH_TRACING, and the command prints that too - but it is on your machine, one environment variable from live. Concept 7 in the README is about what that means.

measure · python
$ ./run -l 8 measure
Try it

Confirm it with the test

An offline test pins every claim on this page: the loop's exact attempt and retrieval counts, the 4/8 → 8/8 result, that the grader cannot see the expected answer, that the graph and the while loop are indistinguishable, that the cycle really cycles, that a runaway graph raises instead of hanging, that the interrupt pauses before generating, that resuming does not re-retrieve, and that a veto never calls the model at all. No network, no model, no API key - and it passes whether or not LangGraph is installed.

test · python
$ ./run -l 8 test
Command

Watch the loop iterate - no code editing

web · python
$ python -m pip install -q -r requirements.txt
web · python
$ ./run -l 8
Experiment

Try - turn off the rewriter

Untick Rewrite the query between attempts and the loop still runs, but it re-searches the same words and gets the same chunks back, so it gives up. A cycle that cannot change its own input is not a cycle - it is a chain that runs slower. Turn it back on and the same question gets answered.

Is 5GHz supported?
Experiment

Try - pause it, then resume it

Leave Human review on 0 pause. The graph stops before generating and shows you what it wants to answer from. Now move it to 1 approve and watch retrievals when paused and retrievals after resume - they are the same number. The graph picked up where it stopped instead of starting the question again. That is what a checkpoint is, and it is the one thing on this page a while loop has no answer for.

How do I wipe the device and start over?
Going further

From demo to production

Pin langgraph exactly - the interrupt API moved recently, and Command(resume=) is not what the tutorials from a year ago show you.

Give the checkpointer real storage, and a retention policy. A checkpoint is a copy of your users' questions and your documents' text, sitting in a database. Decide how long that lives before you write the first one, not after.

Make thread_id a real identity - a conversation, a ticket, a session. A hash of the query means two people asking the same thing share memory.

Set both caps. max_attempts is your domain logic and produces a good answer; recursion_limit is LangGraph's structural floor and raises. Alert on the abstain rate, not on individual abstentions - one refusal is correct behaviour, a rising rate is a corpus problem.

Log every Grade, including its reason. That is what turns the LLM grader's run-to-run spread from a mystery into a metric.

Fold it into Lesson 5. Put all three arms in the golden set and let the evaluation gate tell you whether the loop actually helped on your documents, rather than trusting nine questions someone else wrote.

Recap

What you learned

You turned a linear RAG pipeline into an agent that grades its own retrieval, rewrites its own query, and knows when to stop. It fixed four of the nine questions the chain got wrong and refused the one nothing could answer, for 67% more retrieval calls - one of which bought nothing at all.

Then you priced the graph honestly. The corrective loop is 63 lines of while and zero packages, and it delivered the entire accuracy gain. Swapping those 63 lines for LangGraph cost 172 lines of wiring and 23 packages, for answers that are identical on every question. What it bought is real but narrow: state that outlives the process, a run you can pause and hand to a human, and a topology you can query at runtime.

A framework earns its place when you need what it does, not when it does what you already did. If you never checkpoint and no human is ever in the path, you have bought a state machine you will not use. If they are - and for anything touching an irreversible action, they should be - there is no cheap way to write this yourself.

Next: Lesson 9 · Ollama + function calling - stop routing tools yourself and let a local model decide when to call them, fully offline.

Use ← → arrow keys, the dots, or the buttons. Deep-link a step with #step-N.