Turn the linear RAG pipeline into an agent that grades its own retrieval, rewrites its query and retries - then price the graph honestly against the same loop written as sixty-three lines of while. Three arms over Lesson 7's corpus: the chain fixes 4 of 8, the loop fixes 8 of 8 and correctly refuses the ninth, and the graph matches the loop exactly. Everything LangGraph costs buys checkpoints, interrupts and an inspectable topology - not better answers. Python and Node.js.
New to LangGraph? It is not LangChain 2.0, and it is not an agent framework in the "hand it tools and hope" sense. It is a small state-machine library with checkpoints: you declare a typed state, write nodes that return updates to it, and connect them with edges - including edges that point backwards. It ships Python and JavaScript SDKs, both used here, and it depends on langchain-core, so this lesson sits on top of Lesson 7's dependency rather than beside it.
Lesson 7 rebuilt the pipeline on a framework and asked what it cost. This lesson asks the harder version of that question. A corrective loop - retrieve, grade the evidence, rewrite the query, try again - fixes four of the questions here that a single retrieval gets wrong, and correctly refuses a fifth that nothing in the corpus answers. But you do not need LangGraph for a loop. You need while.
So this lesson runs three arms, not two: the linear chain, the same loop written as sixty-three lines of while, and the loop as a StateGraph. The while loop and the graph agree on every single question. That zero is the lesson. Pick a language above and press → to begin.
A chain runs once. A graph decides whether to run again.
┌───────────────────────────────┐
│ START question, attempt=1 │
└───────────────┬───────────────┘
▼
┌───────────────────────────────┐
┌────────▶│ retrieve BM25 · top-k │ retrievals +1
│ └───────────────┬───────────────┘
│ ▼
│ ┌───────────────────────────────┐
│ │ grade coverage or LLM │ grades +1
│ └──────┬─────────────────┬──────┘
│ weak, and │ │ grounded
│ attempts left│ │
│ ▼ ▼
│ ┌──────────────────────┐ ┌──────────────────────────┐
└──┤ rewrite glossary │ │ review ⏸ interrupt()│
└──────────┬───────────┘ └───┬──────────┬───────────┘
│ approve │ │ veto
edit └───────────────────┤ │
▼ ▼
weak, no attempts left ┌──────────────┐ ┌──────────────┐
─────────────────────────────────▶│ abstain │ │ generate │
└──────┬───────┘ └──────┬───────┘
▼ ▼
END END
Three things this picture has that a chain cannot draw:
◀── an edge that points BACKWARDS (rewrite → retrieve)
⏸ a node that STOPS and waits for a human (review)
▼ ▼ two different ways to finish (generate · abstain)
Measured over Lesson 7's corpus, nine questions:
linear loop (while) graph (LangGraph)
top source correct 4/8 8/8 8/8
correctly abstained 0/1 1/1 1/1
retrieval calls 9 15 (+67%) 15
The while loop and the graph agree on every question. That is not a
disappointment - it is the finding. Everything LangGraph costs is paid
for the three things the loop cannot do at all: keep state past the
process, stop halfway and ask a human, and tell you its own shape.This is the second lesson that installs something, and like Lesson 7 the dependency is part of the argument. ./run -l 8 puts langgraph into the course venv on first use. If you skip it, the demo still runs the linear and loop arms, prints the whole accuracy result, marks the graph column not installed, and exits 0 - which is a better degradation than Lesson 7 managed, because here the dependency is not what produces the result. Run everything from the repo root:
$ ./run -l 8 # Python: the playground (default)
./run -l 8 demo # Python: the three-arm comparison and the scorecard
./run -l 8 --lang node demo # Node.js: same control flow, a different billThree are carried straight from Lesson 7 as controls. Four are phrased the way a real person asks - orange light when the docs say amber status ring, wipe when they say factory reset, RMA when they say warranty claim, 5GHz when they spell out five gigahertz. One is already correct on the first try and makes the loop spend a retrieval for nothing. One has no answer in the corpus at all.
Every question carries its own why, and the demo prints it, so the reason a question exists lives in the data rather than only in this prose.
{
"chunk_size": 700,
"chunk_overlap": 120,
"top_k": 3,
"grade_threshold": 0.67,
"max_attempts": 3,
"questions": [
{ "id": "q1", "ask": "What is the factory reset procedure?",
"expect": "manual.md",
"why": "control - carried from Lesson 7, and it must not regress" },
{ "id": "q2", "ask": "Why does the status ring stay amber?",
"expect": "faq.md",
"why": "control - carried from Lesson 7" },
{ "id": "q3", "ask": "Which endpoint exports the logging buffer?",
"expect": "api.md",
"why": "control - carried from Lesson 7" },
{ "id": "q4", "ask": "The light is orange and it will not connect.",
"expect": "faq.md",
"why": "vocabulary - the reader says 'orange light', the docs say 'amber status ring'" },
{ "id": "q5", "ask": "How do I wipe the device and start over?",
"expect": "manual.md",
"why": "user-speak - the reader says 'wipe', the docs say 'factory reset'" },
{ "id": "q6", "ask": "What paperwork does an RMA need?",
"expect": "warranty.md",
"why": "acronym - the reader says 'RMA', the docs say 'warranty claim'" },
{ "id": "q7", "ask": "Is 5GHz supported?",
"expect": "networking.md",
"why": "notation - the reader types '5GHz', the docs spell out 'five gigahertz'" },
{ "id": "q8", "ask": "Can I mount it sideways?",
"expect": "installation.md",
"why": "FALSE ALARM - already correct on the first try; the grader loops anyway and buys nothing" },
{ "id": "q9", "ask": "What is the MTBF?",
"expect": null,
"why": "not in the corpus at all - the loop has to give up rather than spin" }
]
}Term coverage: what fraction of the query's content words appear anywhere in the retrieved text? It is a real information-retrieval signal, it detects precisely the failure this lesson is about, and missing - the terms that did not land - is exactly what the rewriter needs next.
class CoverageGrader:
"""Term coverage: what fraction of the query's content words does the evidence contain?
This is a real information-retrieval signal, not a stand-in for one. It
detects the exact failure this lesson is about - a reader asking in words the
documents never use - because those words cannot appear in any retrieved
chunk, whatever the retriever ranks first.
It also explains itself. `missing` is the list of terms that did not land,
which is precisely what the rewriter needs, so the two nodes are coupled
through data rather than through a shared hard-coded table.
"""
name = "coverage"
def __init__(self, threshold: float) -> None:
self.threshold = threshold
def grade(self, question: str, query: str, docs: List[rag_core.Chunk]) -> Grade:
terms = rag_core.content_terms(query)
if not terms:
# A query with no content words cannot be judged this way. Say so
# rather than dividing by zero or silently passing.
return Grade(verdict="weak", score=0.0, missing=[],
reason="no content terms to look for", grader=self.name)
haystack = set(rag_core.tokens(rag_core.evidence_text(docs)))
missing = [t for t in terms if t not in haystack]
score = (len(terms) - len(missing)) / len(terms)
grounded = score >= self.threshold
return Grade(
verdict="grounded" if grounded else "weak",
score=round(score, 2),
missing=missing,
reason=(f"evidence covers {len(terms) - len(missing)}/{len(terms)} query terms"
+ ("" if grounded else f"; missing {missing}")),
grader=self.name,
)expect lives in the data and is used only to score the run afterwards. A test asserts that on the function signature, because a grader that can read the label turns the whole measurement into theatre.When the grader says the evidence is weak, the query has to change or the next attempt just repeats the last one. This expands the missing terms into the words the documents actually use, and drops the ones that landed nowhere.
class GlossaryRewriter:
"""Expand the terms the evidence did not cover into the words the docs use.
Only `missing` terms are expanded. Rewriting on a grade with nothing missing
is a no-op, which keeps the rewriter honest: it reacts to the grader's finding
instead of reaching for the table whenever it is called.
"""
name = "glossary"
def __init__(self, glossary: Optional[Dict[str, List[str]]] = None) -> None:
self.glossary = glossary if glossary is not None else rag_core.load_glossary()
def rewrite(self, question: str, query: str, missing: List[str]) -> str:
if not missing:
return query
expanded: List[str] = []
for term in missing:
expanded.extend(self.glossary.get(term, []))
if not expanded:
# Nothing in the table matches. Returning the query unchanged is the
# honest move: the loop will grade it weak again and, at the cap,
# abstain - which is the correct outcome for a question the corpus
# does not answer. Inventing a rewrite here would turn "we have no
# document about this" into "we tried harder", which is worse.
return query
# Keep the terms that DID land, drop the ones that did not, add the
# documents' words. Dropping the misses is the point: "orange" is not in
# any chunk, so leaving it in only dilutes the BM25 score of the rest.
kept = [t for t in rag_core.content_terms(query) if t not in missing]
out: List[str] = []
for term in kept + expanded:
if term not in out:
out.append(term)
return " ".join(out)orange only because somebody wrote orange into a JSON file, and it will not fix the phrasing nobody anticipated. It exists so this page prints the same thing on every machine. That is the entire argument for --llm-rewrite, and the lesson would rather say so than let you discover it.Before any graph: here is the corrective loop as a plain while. It calls the same retrieve, the same grader and the same rewriter the graph will call. Note the early exit - if the rewriter hands back the query unchanged, there is nothing left to try, and a cycle that cannot change its own input is an infinite loop with extra steps.
def run(retriever, question: str, *, grader: Grader, rewriter: Rewriter,
top_k: int, max_attempts: int,
generate: Optional[Callable] = None) -> dict:
"""Retrieve, grade, and re-search with a better query until the evidence holds up."""
query = question
attempt = 1
retrievals = grades = rewrites = 0
trace: List[str] = []
verdicts: List[str] = []
docs: List[rag_core.Chunk] = []
grade = None
while True:
docs = rag_core.retrieve(retriever, query, top_k)
retrievals += 1
trace.append(f"retrieve attempt={attempt} query={query!r} -> {rag_core.sources(docs)}")
grade = grader.grade(question, query, docs)
grades += 1
verdicts.append(grade["verdict"])
trace.append(f"grade {grade['verdict']} score={grade['score']} - {grade['reason']}")
if grade["verdict"] == "grounded":
break
if attempt >= max_attempts:
trace.append(f"abstain {attempt} attempts used, evidence still weak")
break
new_query = rewriter.rewrite(question, query, grade["missing"])
rewrites += 1
if new_query == query:
# The rewriter had nothing left to try. Retrying an identical query
# would retrieve identical chunks and grade identically - a cycle that
# cannot change its own input is an infinite loop with extra steps.
# Stopping here reaches the same decision the cap would, one wasted
# retrieval sooner.
trace.append("abstain rewrite changed nothing - no point asking again")
break
# A rewriter that quietly fell back to the glossary must not look like one
# that produced this query itself.
note = " [fell back to the glossary]" if getattr(rewriter, "fell_back", False) else ""
trace.append(
f"rewrite {query!r} -> {new_query!r} (missing {grade['missing']}){note}")
query = new_query
attempt += 1
grounded = grade["verdict"] == "grounded"
answer = ""
if grounded and generate is not None:
answer = generate(question, docs)
elif not grounded:
answer = rag_core.ABSTAIN_TEXT
return {
"question": question, "query": query, "attempts": attempt,
"retrievals": retrievals, "grades": grades, "rewrites": rewrites,
"sources": rag_core.sources(docs) if grounded else [],
"verdicts": verdicts, "grade": grade, "docs": docs,
"status": "answered" if grounded else "abstained",
"answer": answer, "trace": trace,
}Look at the two kinds of field in one struct. trace, turns and the three counters are Annotated[..., operator.add], so a node returning {"retrievals": 1} adds one to the running total. Everything else is last-write-wins, so a node returning {"query": ...} replaces it.
class AgentState(TypedDict, total=False):
"""Everything that flows between nodes.
Note the two kinds of field, deliberately side by side:
`trace`, `turns`, and the three counters are Annotated with operator.add,
so a node returning {"retrievals": 1} ADDS ONE to the running total.
everything else is last-write-wins, so a node returning {"query": "..."}
REPLACES the query.
That difference is what a reducer is, and it is easier to see in one struct
than in any amount of prose.
"""
question: str # the human's words. Never mutated.
query: str # the CURRENT search query. `rewrite` replaces this.
attempt: int
docs: List[rag_core.Chunk]
sources: List[str]
grade: Grade
trace: Annotated[List[str], operator.add]
turns: Annotated[List[dict], operator.add]
retrievals: Annotated[int, operator.add]
grades: Annotated[int, operator.add]
rewrites: Annotated[int, operator.add]
decision: str # "approve" | "veto" | "edit:<new query>"
answer: str
status: str # "answered" | "abstained" | "vetoed"add_conditional_edges takes a function that returns a string key, and a map from those keys to nodes. rewrite -> retrieve is the cycle. grade has three or four ways out depending on whether a human is in the loop.
def route_after_grade(state: AgentState) -> str:
if state["grade"]["verdict"] == "grounded":
return "review" if human_review else "generate"
if state["attempt"] >= max_attempts:
# The cap is a decision, not a crash. It produces a good answer -
# "not in your documents" - rather than an exception.
return "abstain"
return "rewrite"
def route_after_review(state: AgentState) -> str:
decision = (state.get("decision") or "approve")
if decision == "veto":
return "veto"
if decision.startswith("edit:"):
return "retrieve"
return "generate"
builder = StateGraph(AgentState)
builder.add_node("retrieve", retrieve_node)
builder.add_node("grade", grade_node)
builder.add_node("rewrite", rewrite_node)
builder.add_node("generate", generate_node)
builder.add_node("abstain", abstain_node)
if human_review:
builder.add_node("review", review_node)
builder.add_node("veto", veto_node)
builder.add_edge(START, "retrieve")
builder.add_edge("retrieve", "grade")
targets = {"rewrite": "rewrite", "generate": "generate", "abstain": "abstain"}
if human_review:
targets["review"] = "review"
builder.add_conditional_edges("grade", route_after_grade, targets)
builder.add_conditional_edges("rewrite", _rewrite_router,
{"retrieve": "retrieve", "abstain": "abstain"})
if human_review:
builder.add_conditional_edges("review", route_after_review,
{"generate": "generate", "retrieve": "retrieve",
"veto": "veto"})
builder.add_edge("veto", END)
builder.add_edge("generate", END)
builder.add_edge("abstain", END)
graph = builder.compile(checkpointer=checkpointer)
graph._retrieve_node = retrieve_node # so run() can attach the retriever
return graphget_graph().edges print the back-edge as a value. A while loop cannot answer "what are your edges?" at runtime. Its control flow is if and while; this one is data.The graph stops here and hands back a review packet: the question, the query it ended up searching, how many attempts it took, the citations, and the evidence itself. You resume with Command(resume="approve"), "veto", or "edit:<a better query>" - and that last one re-enters the cycle at retrieve.
def review_node(state: AgentState) -> dict:
"""Stop, and hand a human everything they need to judge the answer.
The payload IS the review packet. That is why this uses the dynamic
`interrupt()` rather than `interrupt_before=["generate"]`: the static form
pauses without telling the reviewer anything, and a human approving a
grounded answer needs to see the evidence it is grounded in.
"""
decision = interrupt({
"question": state["question"],
"query_used": state["query"],
"attempts": state["attempt"],
"citations": state["sources"],
"evidence": [d["text"][:200] for d in state["docs"]],
"ask": "approve | veto | edit:<new query>",
})
out = {"decision": decision, "trace": [f"review human said {decision!r}"]}
if isinstance(decision, str) and decision.startswith("edit:"):
# The reviewer did not just approve or refuse - they handed back a
# better query. That re-enters the cycle at `retrieve`, which is a
# second way into the loop and the reason `review` earns a node
# rather than being a flag on `generate`.
edited = decision[len("edit:"):].strip()
if edited:
out["query"] = edited
out["attempt"] = state["attempt"] + 1
out["trace"] = [f"review human rewrote the query -> {edited!r}"]
return outinterrupt() rather than the older interrupt_before=["generate"], because the payload IS the point. A static pause tells the reviewer nothing, and a human approving a grounded answer needs to see what it is grounded in. The edit branch is also why review earns a node instead of being a flag on generate: it is a second way into the loop.$ python -m pip install -q -r requirements.txt$ ./run -l 8 demo$ npm --prefix node install --silent --no-audit --no-fund$ ./run -l 8 --lang node demoThe headline: 4/8 becomes 8/8, plus the abstention the chain never makes, for 67% more retrieval calls. The four the loop fixes are exactly the four where the reader's words are not the documents' words. q8 was already right and looped anyway - a retrieval that bought nothing, left in because pretending the loop is free would be a lie. q9 has no answer in the corpus, and the loop refuses rather than spinning.
Then the part that matters. The while loop and the graph agree on all nine questions - exactly. LangGraph did not make the agent smarter. Everything it cost bought the three rows the loop leaves blank: state that survives the process, a run you can pause and resume, and a topology you can print. The line counts and package counts are read off disk at run time, so they cannot drift from the code.
Lesson 8 · A stateful agent with LangGraph
corpus: lessons/07-langchain-rag/data/corpus (7 documents - Lesson 7's, by path)
settings: top_k=3 threshold=0.67 max_attempts=3 grader=coverage rewriter=glossary
Nine questions. Three arms. The same corpus and the same settings.
linear loop graph att question
------------------------------------------
q1 PASS PASS PASS 1 What is the factory reset procedure?
q2 PASS PASS PASS 1 Why does the status ring stay amber?
q3 PASS PASS PASS 1 Which endpoint exports the logging buffer?
q4 FAIL PASS PASS 2 The light is orange and it will not connect.
q5 FAIL PASS PASS 2 How do I wipe the device and start over?
q6 FAIL PASS PASS 2 What paperwork does an RMA need?
q7 FAIL PASS PASS 2 Is 5GHz supported?
q8 PASS PASS PASS 2 Can I mount it sideways?
q9 FAIL PASS PASS 2 What is the MTBF?
why each question is here
q1 control - carried from Lesson 7, and it must not regress
q2 control - carried from Lesson 7
q3 control - carried from Lesson 7
q4 vocabulary - the reader says 'orange light', the docs say 'amber status ring'
q5 user-speak - the reader says 'wipe', the docs say 'factory reset'
q6 acronym - the reader says 'RMA', the docs say 'warranty claim'
q7 notation - the reader types '5GHz', the docs spell out 'five gigahertz'
q8 FALSE ALARM - already correct on the first try; the grader loops anyway and buys nothing
q9 not in the corpus at all - the loop has to give up rather than spin
The scorecard
linear (L7) loop (while) graph (LangGraph)
top source correct 4/8 8/8 8/8
correctly abstained 0/1 1/1 1/1
--------------------------------------------------------------------------------
retrieval calls 9 15 (+67%) 15
grade calls 0 15 15
rewrite calls 0 7 7
--------------------------------------------------------------------------------
state survives a process - - checkpointer
pause / resume mid-run - - interrupt()
topology you can query - - get_graph()
The loop fixed 4 questions the chain got wrong (q4, q5, q6, q7), and
refused the 1 it could not answer at all - for 67% more retrieval calls.
q8 was already right on the first try and looped anyway: a retrieval that bought nothing.
The while loop and the graph agree on all 9 questions: exactly.
LangGraph did not make the agent smarter. It made it resumable,
pausable and inspectable - the three rows above that only it can fill.
The topology, as data
nodes ['abstain', 'generate', 'grade', 'retrieve', 'review', 'rewrite', 'veto']
edges
__start__ -> retrieve
abstain -> __end__
generate -> __end__
grade -> abstain
grade -> generate
grade -> review
grade -> rewrite
retrieve -> grade
review -> generate
review -> retrieve
review -> veto
rewrite -> abstain
rewrite -> retrieve <- the cycle
veto -> __end__
A while loop cannot answer the question 'what are your edges?' at runtime.
Its control flow is `if` and `while`; this one is a value you can print.
Human in the loop
question: 'How do I wipe the device and start over?'
This question matters: the manual says a factory reset destroys the logging
buffer permanently, and the warranty page says a claim without that export
cannot be assessed. Answering helpfully and immediately costs the reader
their evidence. So the graph stops and asks.
invoke #1 -> the graph STOPPED. It did not answer.
status paused
get_state().next ('review',)
answer in state? False
review packet citations=['manual.md:1', 'troubleshooting.md:1']
attempts=2 ask='approve | veto | edit:<new query>'
retrievals so far 2
invoke #2 -> Command(resume='approve') on the same thread_id
retrievals now 2 <- unchanged. It resumed; it did not restart.
status answered
invoke #2' -> Command(resume='veto') on a fresh thread
status vetoed - generate never ran
invoke #2'' -> Command(resume='edit:...') - the reviewer hands back a query
retrievals now 3 <- it re-entered the cycle at retrieve
a second way into the loop, which is why review is a node
Memory - two turns on one thread
thread 'aurora' turn 1 What is the factory reset procedure?
cited ['manual.md:1', 'warranty.md:1']
thread 'aurora' turn 2 Which endpoint exports the logging buffer?
cited ['api.md:1', 'manual.md:1']
thread 'somebody-else' turn 1 <- a different thread_id remembers nothing
checkpoints written on 'aurora': 10
MemorySaver keeps this in the process. For state that outlives the process:
pip install langgraph-checkpoint-sqlite
./run -l 8 chat --sqlite /tmp/aurora.db --thread aurora "..."
It is a separate package on purpose - see the bill below.
What it cost
linear loop (while) graph (LangGraph)
code you maintain 68 lines 298 lines 470 lines
requirements lines 0 0 1
packages installed 48 48 71
install size ~34 MB ~34 MB ~43 MB
The corrective loop cost 230 lines and ZERO packages, and it delivered
the entire accuracy gain in the scorecard above. Most of those lines are the
grader and the rewriter, which the graph needs too. The loop ITSELF is
63 lines of `while`.
Swapping those 63 lines for LangGraph cost 172 lines of graph
wiring and 23 packages - for answers that are identical on all nine
questions. What it bought is the three rows the loop could not fill:
checkpoints, interrupts, and a topology you can query.
If you already did Lesson 7 it is only 5 more packages (67 -> 72),
because LangGraph depends on langchain-core and you are already paying for it.
Which of those two numbers applies to you is the whole question.
grader: coverage (deterministic). Run with --llm-grade for the grader you
should actually use; it will not print the same thing twice, and that is the point.Nothing in the corpus says orange, and nothing says light. The documents say amber status ring. The first search returns warranty.md and networking.md - plausible, and wrong. Watch the grader score it 0.00, name what is missing, and watch the second attempt land on faq.md.
The light is orange and it will not connect.The rewriter does its job correctly - it expands the acronym to mean time between failures - and the search still fails, because the corpus genuinely does not cover it. That is the distinction worth having: a bad query and a missing document are different problems, and only one of them is worth retrying. The agent abstains.
What is the MTBF?Every retrieve, every grade, every rewrite, in order. Add --recursion-limit 4 to a question that loops and watch LangGraph's own cap fire instead of yours - they are not the same thing, and the lesson explains why that distinction matters more than it looks.
$ ./run -l 8 trace "What paperwork does an RMA need?"Several turns on the same thread_id, so the checkpointer has something to remember. With no questions it asks two of the demo's; add --sqlite FILE and run it twice to watch the state survive the process.
$ ./run -l 8 chat --thread aurora "What is the factory reset procedure?"This pauses mid-graph on real stdin and shows you the review packet before anything is generated. The question is not chosen at random: the manual says a factory reset destroys the logging buffer permanently, and the warranty page says a claim without that export cannot be assessed. Answering promptly and helpfully costs the reader their evidence. Approve it, veto it, or hand back a better query.
$ ./run -l 8 review "How do I wipe the device and start over?"Run this one. It calls the LLM grader several times against byte-identical evidence and prints the spread of verdicts it gives back, next to the deterministic grader answering the same thing every time.
The deterministic grader is the default here for exactly one reason: demo is committed to a file and diffed by a test, so it has to print the same thing twice. That is a property this lesson needs and your system almost certainly does not. On your own documents, use --llm-grade - it is the better grader, it generalises to phrasings no glossary anticipated, and the fact that it will not repeat itself is something to measure rather than something to avoid.
$ ./run -l 8 spread "Is 5GHz supported?" --runs 7The comparison is deliberately offline. This runs the whole graph with a model at the end - retrieval, grading, rewriting, generation - through the course's configured provider. Add --llm-grade to put the model in the grading node too, which is how you would actually run this.
$ ./run -l 8 ask "Why is the light orange?" --llm-gradeThe topology as data: every node, every edge, and the back-edge marked. This is the thing a while loop cannot answer about itself at runtime, and it costs one call to get_graph().
$ ./run -l 8 graphDo not trust the bill in this lesson - recompute it. This walks LangGraph's actual dependency closure on your machine and counts the code off disk.
Look for langsmith in that list. LangChain's hosted tracing client arrives as a transitive dependency whether you asked for it or not. It is inert until you set LANGSMITH_TRACING, and the command prints that too - but it is on your machine, one environment variable from live. Concept 7 in the README is about what that means.
$ ./run -l 8 measureAn offline test pins every claim on this page: the loop's exact attempt and retrieval counts, the 4/8 → 8/8 result, that the grader cannot see the expected answer, that the graph and the while loop are indistinguishable, that the cycle really cycles, that a runaway graph raises instead of hanging, that the interrupt pauses before generating, that resuming does not re-retrieve, and that a veto never calls the model at all. No network, no model, no API key - and it passes whether or not LangGraph is installed.
$ ./run -l 8 test$ python -m pip install -q -r requirements.txt$ ./run -l 8Untick Rewrite the query between attempts and the loop still runs, but it re-searches the same words and gets the same chunks back, so it gives up. A cycle that cannot change its own input is not a cycle - it is a chain that runs slower. Turn it back on and the same question gets answered.
Is 5GHz supported?Leave Human review on 0 pause. The graph stops before generating and shows you what it wants to answer from. Now move it to 1 approve and watch retrievals when paused and retrievals after resume - they are the same number. The graph picked up where it stopped instead of starting the question again. That is what a checkpoint is, and it is the one thing on this page a while loop has no answer for.
How do I wipe the device and start over?Pin langgraph exactly - the interrupt API moved recently, and Command(resume=) is not what the tutorials from a year ago show you.
Give the checkpointer real storage, and a retention policy. A checkpoint is a copy of your users' questions and your documents' text, sitting in a database. Decide how long that lives before you write the first one, not after.
Make thread_id a real identity - a conversation, a ticket, a session. A hash of the query means two people asking the same thing share memory.
Set both caps. max_attempts is your domain logic and produces a good answer; recursion_limit is LangGraph's structural floor and raises. Alert on the abstain rate, not on individual abstentions - one refusal is correct behaviour, a rising rate is a corpus problem.
Log every Grade, including its reason. That is what turns the LLM grader's run-to-run spread from a mystery into a metric.
Fold it into Lesson 5. Put all three arms in the golden set and let the evaluation gate tell you whether the loop actually helped on your documents, rather than trusting nine questions someone else wrote.
You turned a linear RAG pipeline into an agent that grades its own retrieval, rewrites its own query, and knows when to stop. It fixed four of the nine questions the chain got wrong and refused the one nothing could answer, for 67% more retrieval calls - one of which bought nothing at all.
Then you priced the graph honestly. The corrective loop is 63 lines of while and zero packages, and it delivered the entire accuracy gain. Swapping those 63 lines for LangGraph cost 172 lines of wiring and 23 packages, for answers that are identical on every question. What it bought is real but narrow: state that outlives the process, a run you can pause and hand to a human, and a topology you can query at runtime.
A framework earns its place when you need what it does, not when it does what you already did. If you never checkpoint and no human is ever in the path, you have bought a state machine you will not use. If they are - and for anything touching an irreversible action, they should be - there is no cheap way to write this yourself.
Next: Lesson 9 · Ollama + function calling - stop routing tools yourself and let a local model decide when to call them, fully offline.
#step-N.