A single AI agent in a loop for one pass is subject to specific limitations. It performs within one context window, executes its loops in sequence, and often performs self-validation on its outputs. Graph engineering is the process of designing the graph on which the AI agent will perform its work, which means constructing the nodes responsible for carrying out the work, the edges responsible for directing the flow of tasks from one node to another, and the shared state that will transport information between the nodes in the graph. The concept began to gain traction in mid-2026 as engineers began to integrate multiple agents specialized in their fields of operation as opposed to using a single agent for a loop.
What Will I Learn?
What Is Graph Engineering?
The graph engineering methodology comprises the design of three elements of an agent system in AI agents: Nodes, Edges and Shared State. Nodes represent computation elements. Edges determine which node will run next. Shared State is the representation of the data structure used by all the nodes in the process of computation.
The node is not necessarily a model call. Some examples of nodes include:
- LLM call node that generates or scores the input text
- Code that is deterministic in nature and acts on the input
- Router that identifies the input and branches
- Verifier that verifies the input against some criteria
- Approval by a human being that halts the run
The types of edges can also vary. A fixed edge is an edge that is responsible for transferring the run from a specific node to another specific node. Conditional edge is one that evaluates the current state and chooses the next node among various choices available. Fan-out edge is used to initiate several nodes simultaneously from a single node. Fan-in edge combines the results of several nodes initiated in parallel from a single node.
Where the Term Originated
The name “Graph Engineering” emerged when discussing AI-engineering on the social network X in mid-July 2026 and is the layer above the “loop engineering.”
However, technically, these concepts are quite old. The modeling of complicated workflows in such a way has been done by state machines, DAGs, and workflow orchestration engines for many years, even before the existence of large language models. Temporal and Apache Airflow are only two of those systems, which run graph-shaped stateful workflows in a non-AI context from the 2010s. In December 2024, Anthropic released a framework named “Building Effective Agents,” which already described many current topologies in graph engineering more than a year before the term became popular.
Artificial Intelligence (AI) Course
Average time: 4 month(s)
Skills you’ll build: Python for AI, Machine Learning, Neural Networks, NLP Basics, AI Tools (ChatGPT, Copilot)
Graph Engineering’s Position in the AI Engineering Stack
The definition of graph engineering is that it is the outermost of the five layers of an AI application, where each successive layer is more distant from the model.
| Layer | What it engineers | Core design question |
|---|---|---|
| Prompt | The single request sent to the model | Is the request clear and complete? |
| Context | What information the model receives | Does the model have the data it needs? |
| Harness | Tools, memory, and scaffolding around the model | Can the model act on its environment and retain information? |
| Loop | The repeat cycle one agent runs | When does the agent check its work and stop? |
| Graph | Coordination between multiple agents or steps | Which node runs next, and what state does it receive? |
Every stage depends on the functioning of the stage below it. A graph that has its loops poorly constructed becomes a system that fails at many points, rather than few, because a defect in a node within a graph is a failure in the same manner as a defect in an individual loop since the graph is just an addition to the reliability of the nodes.
Graph Engineering vs. Loop Engineering
Loop Engineering and Graph Engineering involve designs of various entities. In Loop Engineering, the design is based on the design of a single iteration which involves an entity that thinks, acts, senses the results of the action, and determines whether it is complete. In Graph Engineering, the design involves how several iterations work together with one another.
| Dimension | Loop Engineering | Graph Engineering |
|---|---|---|
| Unit of design | One iteration cycle: act, observe, verify | The topology: which nodes exist and which transitions are permitted |
| Who selects the next step | The model, at runtime | Code, using rules the engineer defines in advance |
| Path on repeated runs | Can differ each run | Consistent for a given input and state |
| Primary failure mode | An unbounded loop that spends tokens without finishing | Coordination failures: conflicting state writes, error propagation across branches |
| Best suited for | A single, well-scoped task with a clear stopping condition | Work that splits into distinct specialties with defined hand-offs |
The Three Building Blocks: Nodes, Edges, and State
Nodes
The node does precisely one piece of work and produces a new state as output. According to LangGraph’s documentation, the node can have either a model call or any arbitrary code with no need for a model to be involved. The factor that has the greatest impact on the reliability of the system is figuring out whether the node needs a model call or can be written as a deterministic function.
Edges
The edge determines which node will execute next and how static or dynamic that decision is. The difference that makes the most impact on the operation of the system is the difference of control after the hand-off. In the case of the tool metaphor, the manager node invokes the specialist node to execute a sub-task of limited scope while retaining control of the entire execution. In the hand-off case, the specialist node executes the remaining execution after taking control from the manager node.
Shared State
The shared state is the data structure accessed by all the nodes of the graph. The data structure generally comprises the original task to be performed, its intermediate output, and any verdict reached on the basis of the verification performed by the verifier nodes. Any state needs a defined merger policy in case there could be any conflict due to parallel writing of the same field by two different nodes, otherwise one of them simply overrides the other.
Why Teams Move From a Single Loop to a Graph
Four drawbacks of the single-agent loop approach make the switch towards the graph approach necessary.
- Context exhaustion. The longer the research or migration procedure, the more intermediary information there is for the procedure. The degradation of the quality of output happens prior to the context limitation of the model being reached.
- Serialized Latency. A loop which processes ten independent sub-procedures one after another would take about ten times as much time as parallel processing.
- Missing Isolation. If ten unrelated sub-procedures are processed in one context, a failed retrieval in one procedure might affect the way the model thinks about an unrelated sub-procedure.
- Weak self-grading. Self-checking of output by a model is not a perfect judge, and there is no other node within the single loop to act as a checker.
The engineering account written by Anthropic in June 2025 regarding their multi-agent research system describes how a multi-agent design performed better than an equivalent single-agent design in their internal research assessment. According to the document, the multi-agent design worked 90.2% better in comparison to a single-agent design. Also, Anthropic says that multi-agent design used 15 times more tokens than a single chat interaction. However, Anthropic’s own assessment shows that it is worth it only when it comes to tasks which are valuable enough to justify the increased token usage, but recommends not using the design in cases where the task is such that all the agents need the same context or the subtasks depend heavily on each other – most coding tasks are in this list.
There is a single property that determines whether or not the extra cost justifies the task: the independence of the subtasks. A research task that is divided into five independent sub-questions enables the subtask to be evaluated independently. But in the case of coding tasks where the task is split into five agents, the subtasks cannot be evaluated independently since any change to one file implies knowledge of any change to another file.
Artificial Intelligence (AI) Course
Average time: 4 month(s)
Skills you’ll build: Python for AI, Machine Learning, Neural Networks, NLP Basics, AI Tools (ChatGPT, Copilot)
Common Agent Graph Topologies
Production graph systems are built from a small number of recurring topologies.
Sequential Pipeline
In a sequential pipeline topology, the task is broken up into a sequence of steps, and each step consumes the output from the previous step. In between steps, programmatic tests take place. The advantage of this topology is that it is predictable and cheap to debug since the code, not a model, determines the sequence.
Routing
Routing classifies the input and routes it to one of several branches that are specialized for that case. Each of the branches can be individually optimized, rather than a single prompt trying to accommodate all possible cases.
Parallel Fan-Out and Fan-In
The fan-out/fan-in topology divides the task into several independent branches that operate in parallel, then combines the outputs. There are two common variants of this topology: sectioning, where each branch deals with its own part of the task; and voting, where the task is executed several times and the results are checked for agreement.
Orchestrator-Worker
The orchestrator-worker topology uses a lead agent to decompose the task at runtime, delegate subtasks to the workers, and then assemble the results. This is the topology that is used in Anthropic’s research system: a lead agent creates a set of subagents that consider different aspects of the task in separate context windows.
Execution Graphs vs. Knowledge Graphs
An execution graph and a knowledge graph only have the same name and nothing else. An execution graph is a representation of control flow, where nodes are units of work, the edges represent allowed transitions between nodes, and the graph itself lives only during one run of the system. A knowledge graph represents a domain, where nodes represent entities, the edges represent relationships, and the graph remains static between runs as a description of what exists in a business or data set. A system can use one, both, or none of them.
The two intersect only at one very specific point. Any node within an execution graph that makes a query on enterprise data should perform a query, and an execution graph itself cannot verify whether the query is semantically correct. The generated query can be syntactically correct, successfully executed, but still reference an entity that does not exist by the mentioned name or aggregate the data on a wrong level of granularity. Routing decides when a node reads data, while data model determines what a read can return.
Artificial Intelligence (AI) Course
Average time: 4 month(s)
Skills you’ll build: Python for AI, Machine Learning, Neural Networks, NLP Basics, AI Tools (ChatGPT, Copilot)
Build a Real Agent Graph
The example below uses LangGraph, the graph orchestration library maintained by LangChain.
Step 1: Define the shared state schema.
from typing import TypedDict, List
class ResearchState(TypedDict):
question: str
documents: List[str]
answer: str
score: int
attempts: int
Step 2: Define each node as a function that accepts and returns state.
def search(state: ResearchState) -> ResearchState:
state["documents"] = retriever.search(state["question"])
return state
def write(state: ResearchState) -> ResearchState:
state["answer"] = model.generate_answer(state["question"], state["documents"])
return state
def review(state: ResearchState) -> ResearchState:
state["score"] = model.rate_answer(state["answer"], state["documents"])
state["attempts"] += 1
return state
Step 3: Build the graph and register a conditional edge.
from langgraph.graph import StateGraph, START, END
def route_after_review(state: ResearchState) -> str:
if state["score"] >= 8:
return END
if state["attempts"] >= 3:
return END
return "write"
graph = StateGraph(ResearchState)
graph.add_node("search", search)
graph.add_node("write", write)
graph.add_node("review", review)
graph.add_edge(START, "search")
graph.add_edge("search", "write")
graph.add_edge("write", "review")
graph.add_conditional_edges("review", route_after_review)
app = graph.compile()
Step 4: Run the graph.
result = app.invoke({
"question": "What is our refund policy?",
"documents": [],
"answer": "",
"score": 0,
"attempts": 0,
})
print(result["answer"])
The routing decision in route_after_review runs in code, not inside the model. The model produces a score; the graph’s own logic decides whether that score ends the run or sends the task back to the write node. This separation — model output as data, routing as code — is the core mechanical difference between a graph and an open-ended loop.
Graph Engineering Frameworks Compared
Four frameworks account for most production graph-engineering implementations as of mid-2026.
| Framework | Maintainer | Status | Checkpointing | Best suited for |
|---|---|---|---|---|
| LangGraph | LangChain | Actively developed | Built in, via a configurable checkpointer | General-purpose agent graphs requiring fine-grained node control |
| Google ADK (Agent Development Kit) | Actively developed | Built in | Teams already using Google Cloud/Gemini infrastructure | |
| Microsoft Agent Framework | Microsoft | Actively developed; successor to AutoGen | Built in | Teams requiring A2A and MCP protocol support out of the box |
| CrewAI | CrewAI, Inc. | Actively developed | Available in Flows | Role-based multi-agent systems with a lower setup overhead |
LangGraph is described by LangChain as a low-level orchestration runtime that enables the creation of long-lived and stateful agents. The user has to create a StateGraph, register nodes, and register edges in LangGraph, enabling precise control while incurring a higher amount of setup code compared to other frameworks.
Google ADK documentation describes the framework as an orchestrator of complex operations in graph-based architecture. Named sequential, parallel, and loop workflow agents are included in the package, alongside agent routing and Agent2Agent (A2A) protocol implementation.
Microsoft AutoGen, which introduced multi-agent orchestration in the form of GraphFlow package, is currently in maintenance mode. According to Microsoft’s GitHub repository, the project does not get any new features and is managed by the community. Microsoft suggests moving on to Microsoft Agent Framework, which is the named successor of the project and which supports multi-agent orchestration in A2A and MCP protocols.
CrewAI is an open-source Python library for orchestrating agents, divided into “crews,” which are more open to arbitrary collaboration and “flows,” which allow a more graph-like approach to specifying the order in which operations need to be executed.
A2A (Agent2Agent) is an open standard that lets agents written in different frameworks delegate tasks to each other across systems’ boundaries. It is the layer that becomes relevant if the graph nodes belong to different teams or different vendors rather than implemented in a unified code base. Google ADK and Microsoft Agent Framework both support A2A natively, which means a node written in one of them can delegate a task to a node implemented in the other without any integration layer.
The choice of the framework is determined by the graph’s boundary condition more than by any comparison of features. If the entire graph belongs to one team and uses one model provider, LangGraph or CrewAI would make the best choice due to lower integration cost. Otherwise, if graph nodes belong to multiple teams, vendors, or companies, the best choice is the framework that supports A2A.
When Should You Use a Graph?
A single well-scoped task with one clear verifier is a loop, and building a graph for it adds unnecessary overhead. The table below lists the signals that favor each approach.
| Signal | Loop is sufficient | Graph is warranted |
|---|---|---|
| Shape of the task | One job with a defined finish line | Splits into distinct specialties that hand off to each other |
| Parallelism | Steps run in sequence | Independent steps benefit from running at once |
| Tools or models per step | Same tools throughout | Different models or toolsets required per step |
| Control flow | One agent can operate safely without close supervision | Explicit, auditable routing between roles is required |
| Failure isolation | A failed step can simply retry | One bad step must not corrupt the rest of the run |
| Verification | The agent checks its own output | A separate, independent node reviews another node’s output |
A Five node graph that would sum up an entire PDF (fetcher, chunker, summarizer, reviewer, and formatter) will end up making a system that works more slowly and is harder to debug when compared to one agent fetching the document and then writing its summary. In addition, five sources of morning market briefs that fan out in parallel, synthesize the outcome, write down a summary, and then send this summary to an independent reviewer node can do things in a manner which a loop simply cannot.
What a Graph Actually Costs
The cost of the tokens depends on the number of nodes and the number of interactions of the model with each node, rather than on the complexity of the task itself. The figure provided by Anthropic of 15 times the cost of tokens of one chat interaction serves as a starting point for this calculation.
Worked example, using Claude Sonnet 5 API pricing as of August 2026 — $2.00 per million input tokens and $10.00 per million output tokens under introductory pricing through August 31, 2026 :
A single-loop research task using approximately 40,000 input tokens and 10,000 output tokens costs:
- Input: 40,000 ÷ 1,000,000 × $2.00 = $0.08
- Output: 10,000 ÷ 1,000,000 × $10.00 = $0.10
- Total: approximately $0.18 per run
The same task rebuilt as a five-node graph, applying Anthropic’s reported 15x token multiplier and holding the input-to-output ratio constant, uses approximately 600,000 input tokens and 150,000 output tokens:
- Input: 600,000 ÷ 1,000,000 × $2.00 = $1.20
- Output: 150,000 ÷ 1,000,000 × $10.00 = $1.50
- Total: approximately $2.70 per run
The token multiplier is what drives the cost rather than the cost per token. Cost can be reduced by lowering the number of nodes making model calls, or changing a model call node to a deterministic one when possible, than using a cheaper model.
Failure Modes and How to Avoid Them
| Failure mode | Description | Mitigation |
|---|---|---|
| State contention | Parallel branches write to the same state field without a merge rule; results silently disappear | Define an explicit merge rule for every field a parallel branch can write |
| Fan-out cost creep | Each additional parallel worker multiplies token spend regardless of its contribution | Cap the number of parallel branches; review whether each branch is necessary |
| Error propagation | A bad output from one node passes to downstream nodes with no signal that it is incorrect | Give each node an explicit error field in state, and route on it |
| Non-deterministic routing | A run that fails once may take a different path on retry, complicating reproduction | Log the routing decision at every conditional edge, not only node outputs |
| Over-engineering | A graph is built where one loop would have completed the task | Start with the simplest topology that works; add nodes only when a specific limitation requires it |
Among the five failure modes, the most frequently occurring one in practice is that of over-engineering, since while adding an additional node comes with no direct cost in terms of any visible expenditure, like a bug does, the cost of the additional node surfaces only later through slower execution, greater token spend, and increased attack surface.
Human-in-the-Loop Nodes
The human in the loop node stops the graph prior to performing a certain task and restarts it only after receiving confirmation, making changes, or rejecting the current state from a human. Typical actions which get gated through this process include sending external communication, making a refund, deleting a record, and releasing code into a production environment. As all the graph’s progress exists within the shared state object, a human can review and alter the state before proceeding with a run, without the rest of the nodes being aware of the human intervention.
Specifying an approval step as a human-checkpoint node with an input edge turns the condition into a structural constraint that cannot be ignored during the run. Specifying it solely as instructions in a prompt turns the condition into just a suggestion.
Artificial Intelligence (AI) Course
Average time: 4 month(s)
Skills you’ll build: Python for AI, Machine Learning, Neural Networks, NLP Basics, AI Tools (ChatGPT, Copilot)
Testing, Observability, and Governance
Testing a Node in Isolation
Every node in the graph is a function that takes an input and generates an output and therefore it can be tested in an isolated way from the model and the rest of the graph. Testing a node can provide a state object, execute the node and assert its state without the need for a live model call for deterministic nodes or a single controlled model call for non-deterministic nodes.
Tracing a Run
For each graph, there is one trace per node and trace per edge traversal, including parallel branches of execution and where they combined their outputs. Debugging a graph requires knowing how a specific run has executed, which requires logging the actual routing decision at each conditional edge, not just the results of each node’s execution. Specialized tools for this are LangSmith (tracing platform of LangChain) and Langfuse (open-source LLM observability platform). OpenTelemetry instrumentation modified to capture agent traces is another option.
Evaluation of a graph is connected to node boundaries, not the whole system. Every node that has an input and an output can be evaluated in isolation and turns an otherwise opaque multiple step system into something you can measure.
Making Policy Structural
Both cost and policy are transformed into properties of the graph itself and no longer of the system as a whole. The token expenditure is tied to the node itself, meaning that the costly element of the system can be recognized in the individual node itself rather than in a sum total. Anthropic’s multi-agent research framework states that useful evaluation signals can be acquired even with as few as twenty queries if the effect sizes are large enough, using an LLM as a judge.
Security: Preventing One Node From Compromising the Rest
A prompt injection attack working on one particular node in a graph will be able to spread to all subsequent nodes where the output of that particular node is used as trusted input. There are two defenses against such a scenario. Firstly, the permissions for each node should be restricted to just what is necessary for its purpose — the task of a summarization node is different from the task of sending emails, and the task of a research node is not the same as deleting entries. Secondly, the output of any node which deals with untrusted data (websites, uploaded files, or third-party API responses) must be considered as data and not instructions.
Graph Engineering vs. Traditional Workflow Orchestration
Graph manipulation technologies and workflow management systems like Temporal, Apache Airflow, and AWS Step Functions address similar but different challenges. Workflow management systems were developed to manage deterministic, long-running tasks with reliable execution, automatic retry, and durable state that is resilient even in the face of process restarts — features that have been around since before large language models existed. The Temporal platform gives you durable execution; a workflow state is persisted automatically via event sourcing without developer effort.
A model-driven agent graph constructed using a platform like LangGraph or Google ADK incorporates an element that neither of these two platforms inherently includes: nodes whose subsequent action is determined by a model in real time rather than by pre-written code. Indeed, some of the existing systems make use of both approaches — using Temporal for example to ensure durability of execution in the outer workflow while implementing each of the actions within that workflow via model-driven agent nodes.
Best Practices Checklist
- Use the simplest possible topology that works. Only add more nodes if you identify a concrete limitation in the current configuration.
- Make edges explicit wherever possible. Let the model decide only in places where it’s necessary.
- Keep the verifier as a separate node from the producer. A node should not be able to evaluate its own output.
- Type the shared state and give a merge rule for all fields that a parallel branch writes to.
- Set an upper bound on every cycle. The unbounded loop-back edge is the most frequent cause of exponential token spend.
- Every path in the graph must have an endpoint; this includes failure paths.
- Write errors to state and route on them instead of stopping the whole execution with a single failed tool call.
- Record the routing decision on each conditional edge, not just the output of each node.
- Test nodes separately and then test the graph.
- Choose an existing framework over rolling your own orchestration layer unless the requirements of your project are not covered by LangGraph, Google ADK, Microsoft Agent Framework, or CrewAI.
Frequently Asked Questions
Q1. What is graph engineering?
Ans. Graph engineering is the design of nodes, edges, and shared state in the AI agent system rather than depending upon a single agent running in a loop.
Q2. Is graph engineering the same as LangGraph?
Ans. No. LangGraph is a framework that uses the principle of graph engineering. Google ADK, Microsoft Agent Framework, and CrewAI are other frameworks that use the same principle but through different APIs.
Q3. Is graph engineering the same as a knowledge graph?
Ans. No. The knowledge graph is a model of the entities and relations in the domain and is persistent between runs. The execution graph that is designed through graph engineering is only concerned with modeling the control flow of a particular run.
Q4. Do I need a framework, or can a graph be built without one?
Ans. It is possible to create a graph without a framework using simple functions and a routing dictionary as shown above. The frameworks are used by production-level systems because they give functionality for checkpointing, tracing, and handling of parallel execution that otherwise will have to be implemented.
Q5. How much more does an agent graph cost to run than a single agent loop?
Ans. Anthropic estimated its multi-agent research experiment setup cost 15 times as many tokens than a single chat session. How much greater the exact multiplier is going to be for your particular graph will depend on the number of its nodes and their token consumption.
Q6. How is graph engineering different from workflow orchestration tools like Temporal or Airflow?
Ans. The workflow orchestration tools manage deterministic, lengthy, with guarantee of execution, and durable-state processes. In the case of an agent graph, new nodes are added at runtime based on a model’s decision rather than pre-written code.
Q7. Is “graph engineer” a formal job title?
Ans. As of June 2026, no official job title exists. The core skill set involved with this kind of position — development of multiple-node agent systems — can be found in job titles like AI engineer and machine learning engineer.
Q8. How many nodes is too many for one graph?
Ans. There is no universal limit. Your graph needs to be stopped once its further nodes cease to carry out tasks that cannot be executed by a simpler system architecture; graphs with dozens of nodes and untyped state are frequently mentioned as problematic.
Q9. Can a graph call another graph?
Ans. Yes. An element within a graph may call an entirely separate graph as a procedure, which is a typical way of reusing a tested topology, such as a research-verify pattern, in many parent graphs.
Q10. Does adding graph engineering guarantee better output quality?
Ans. No. The graph ensures better cooperation, auditability, and robustness of failures. Quality of output depends on the quality of each individual node loop, and so, a bad graph made up of untested loops results in a system which breaks in more places than just one weak loop.
Q11. What is the fastest way to tell if a project needs a graph before building one?
Ans. Write out a task in the form of one instruction addressed to one agent. If your instruction uses the word “then” more than two times while describing distinct specialties and not just a series of steps necessary for completing one single specialty, you need a graph. If not, a loop will be enough.
Conclusion
The artifact being designed in agent-AI architectures has undergone two changes in as many years; from the exact wording of the prompt to the structure of the loop used to repeat it, to the system topology within which those loops execute. Nodes, edges, and common state define the cost and correctness of a multi-agent system even before its execution begins. The decision prior to all of this regarding whether the problem needs to be solved in a graph, or whether a well-specified loop and verifier could solve the problem far more cheaply, is the most often ignored one.