Agent Orchestration (Part 3)

Neil Haddley • September 24, 2026

Finding a real agent through a2a-registry.org, then orchestrating it from a LangGraph graph with a local Ollama model: discover the Agent Card, delegate over A2A, verify the reply against the official source, and pause for human approval

AIa2alanggraphollamaagent-orchestrationagent-registryfederal-register

This is part 3 of a series on agent orchestration. Part 1 measures four LangChain patterns for agents you build yourself, and part 2 introduces the Agent2Agent protocol.

My previous post, part 2, ran a "hello world" agent I wrote myself, to see the protocol's mechanics with nothing else in the way. This post uses an agent I did not write. I found it in a public registry, a2a-registry.org. In the LangGraph graph I fetch its Agent Card again on every run, so the run checks what the agent advertises today and not what the registry listed when I looked, and then I call it. The graph runs on my own machine with a model hosted by Ollama. The graph is the point: it shows what an orchestrator is responsible for when it depends on an agent it does not control.

The task is a regulatory watch. The graph asks the remote agent for a briefing of new Federal Register documents, checks every row against the official source, has a local model score each verified document for a stated job role, drafts a short note about the relevant ones, and stops for my approval before anything would be sent.

Every output below is from a real run against live services on 20 September 2026. The briefing changes daily, so a run today will differ.

a2a-registry.org

An Agent Card only helps once you know where to fetch it. a2a-registry.org covers the step before that, discovering agents you did not already know about. It describes itself as "the definitive directory for the Agentic Web" and, at the time of writing, lists 301 agents, 74 of them verified. Verification means the owner proved control through a DNS TXT record, a linked GitHub account or a GoDaddy ANS name. Adding a listing needs no account: you paste an agent's URL and the registry reads the Agent Card.

The registry is an A2A agent itself. Its own card declares a JSON-RPC interface with a search_agents skill that takes a natural-language query. It also declares bearer-token authentication, and an unauthenticated call returns HTTP 401 with a pointer to a token sign-up page. I did not create an account, so I browsed the categories by hand instead.

Listings vary a great deal, and the registry's own health status can be stale. Four agents I looked at all showed the identical "14 consecutive health check failures", and a direct request for one of their cards succeeded, which suggests the registry's checker had stalled, not that four agents failed together. Before choosing, I checked each candidate's card, or its registry page where I did not fetch the card, for two things: which transport it declares, and whether it charges.

AgentHow I checkedResult
AgentNative Data ExchangeCard fetched, several calls madeJSON-RPC 1.0, declares no auth, free sample tier. Chosen
CharitySenseCard fetchedTransport is OPENAPI: a REST API described with an A2A-shaped card, not an A2A message exchange
ForgeMesh travel and faresRegistry pageEvery call is paid in USDC through x402 crypto payments
The registry's own agentCard fetched, call returned 401Needs an account token

The agent: AgentNative Data Exchange

The registry listing points at an Agent Card, and the first thing the graph does is fetch it.

The Agent Card

This is the card exactly as the agent serves it at https://agentnative.cazimedia.com/.well-known/agent-card.json. I fetched it with curl and pretty-printed it:

JSON
1{
2    "name": "AgentNative Data Exchange",
3    "description": "Normalized official government and public datasets across federal, state and city agencies, census, police, education and schools, with provenance, aggregations, insights and free samples.",
4    "url": "https://agentnative.cazimedia.com/a2a",
5    "version": "0.3.0",
6    "supportedInterfaces": [
7        {
8            "url": "https://agentnative.cazimedia.com/a2a",
9            "protocolBinding": "JSONRPC",
10            "protocolVersion": "1.0"
11        },
12        {
13            "url": "https://agentnative.cazimedia.com",
14            "protocolBinding": "HTTP+JSON",
15            "protocolVersion": "1.0"
16        }
17    ],
18    "capabilities": {
19        "streaming": false,
20        "pushNotifications": false,
21        "extendedAgentCard": false
22    },
23    "defaultInputModes": [
24        "text/plain",
25        "application/json"
26    ],
27    "defaultOutputModes": [
28        "text/plain",
29        "application/json"
30    ],
31    "skills": [
32        {
33            "id": "public-data-search",
34            "name": "Search official government and public datasets",
35            "description": "Discover federal, state, city, census, police, education and school data with official provenance.",
36            "tags": [
37                "government data",
38                "public datasets",
39                "statistics",
40                "provenance"
41            ],
42            "examples": [
43                "Find official datasets about public schools."
44            ]
45        },
46        {
47            "id": "on-demand-materialization",
48            "name": "Make discovered data query-ready",
49            "description": "Prioritize a compatible official dataset for detached ingestion and poll the durable task through completion.",
50            "tags": [
51                "materialization",
52                "async task",
53                "data import"
54            ],
55            "examples": [
56                "Materialize disc_0123456789abcdef01234567 and tell me when it is query-ready."
57            ]
58        },
59        {
60            "id": "imported-data-sample",
61            "name": "Sample normalized government datasets free",
62            "description": "Inspect official rows, summaries, deterministic insights, freshness and provenance before payment.",
63            "tags": [
64                "free sample",
65                "open data",
66                "normalized records"
67            ],
68            "examples": [
69                "Show me useful free data with provenance."
70            ]
71        },
72        {
73            "id": "imported-data-query",
74            "name": "Query and aggregate official datasets",
75            "description": "Filter, group and aggregate normalized public records across government agencies.",
76            "tags": [
77                "aggregation",
78                "analysis",
79                "official records"
80            ],
81            "examples": [
82                "Group these public records by agency."
83            ]
84        },
85        {
86            "id": "coverage-status",
87            "name": "Inspect source coverage",
88            "description": "Inspect current catalog traversal, materialization, queue, failure, and publication freshness state.",
89            "tags": [
90                "coverage",
91                "freshness",
92                "data sources"
93            ],
94            "examples": [
95                "Which official sources are currently queryable?"
96            ]
97        },
98        {
99            "id": "federal-register-query",
100            "name": "Query Federal Register",
101            "description": "Query normalized rules, proposed rules, notices, and presidential documents with bounded date and type filters.",
102            "tags": [
103                "Federal Register",
104                "regulations",
105                "rules",
106                "notices"
107            ],
108            "examples": [
109                "Give me a free current Federal Register briefing with provenance."
110            ]
111        }
112    ]
113}

What matters in it:

- supportedInterfaces lists two ways to reach the agent: JSONRPC at https://agentnative.cazimedia.com/a2a and HTTP+JSON at the base URL, both at protocol version 1.0. The graph's discover step prints the first entry, JSONRPC 1.0, and the hand-made call below uses the JSON-RPC one.

- capabilities.streaming is false, so a call returns one message and not a stream of updates.

- There is no securitySchemes key, so the card declares no authentication.

- Six skills are listed. The one this post uses is federal-register-query, and discover stops the run if that id is missing. The card gives an example prompt for it and no input schema, so the example sentence is the only guidance on how to ask. That is why the graph sends that exact sentence. A different free-text prompt I tried, a search for datasets about public schools, returned a capabilities menu and no data.

- on-demand-materialization advertises a "durable task" you poll to completion, which would exercise A2A's task states. I did not use it.

Calling it

The a2a-sdk client is given nothing beyond the base URL:

PYTHON
1client = await create_client("https://agentnative.cazimedia.com", ClientConfig(streaming=False))
2message = Message(message_id=str(uuid.uuid4()), role=Role.ROLE_USER,
3                  parts=[Part(text="Give me a free current Federal Register briefing with provenance.")])
4async for event in client.send_message(SendMessageRequest(message=message)):
5    ...

To see the wire format, I made the same call by hand once with curl, using the JSON-RPC interface from the card and the A2A-Version: 1.0 header:

BASH
1curl -X POST https://agentnative.cazimedia.com/a2a \
2  -H "Content-Type: application/json" -H "A2A-Version: 1.0" \
3  -d '{"jsonrpc":"2.0","id":1,"method":"SendMessage","params":{"message":{"role":"ROLE_USER","message_id":"wire-check-1","parts":[{"text":"Give me a free current Federal Register briefing with provenance."}]}}}'

The response is a JSON-RPC envelope holding one A2A message, from the ROLE_AGENT, with a single text part:

JSON
1{"jsonrpc": "2.0", "id": 1, "result": {"message": {"messageId": "...", "role": "ROLE_AGENT", "parts": [{"text": "<a JSON document, shown next>"}]}}}
The reply

The text part is itself a JSON document. Here it is, decoded and abridged: I shortened the long title, abstract and URL strings and collapsed one field, and changed nothing else.

JSON
1{
2  "type": "federal-register-briefing",
3  "status": 200,
4  "data": {
5    "dataset": "us-federal-register-documents",
6    "sample": true,
7    "sample_basis": "Aggregates cover the 25 newest documents; three example records are included.",
8    "rows": [
9      {
10        "document_number": "2026-19251",
11        "title": "Presidential Determination on Major Drug Transit or Major Illicit Drug Producing Countries ...",
12        "publication_date": "2026-09-18",
13        "type": "Presidential Document",
14        "agencies": [
15          "Executive Office of the President"
16        ],
17        "abstract": null,
18        "url": "https://www.federalregister.gov/documents/2026/09/18/2026-19251/presidential-determi ..."
19      },
20      {
21        "document_number": "2026-19222",
22        "title": "Employment in the Excepted Service",
23        "publication_date": "2026-09-18",
24        "type": "Proposed Rule",
25        "agencies": [
26          "Personnel Management Office"
27        ],
28        "abstract": "The Office of Personnel Management (OPM) proposes to amend its regulations governing the e ...",
29        "url": "https://www.federalregister.gov/documents/2026/09/18/2026-19222/employment-in-the-ex ..."
30      },
31      {
32        "document_number": "2026-19211",
33        "title": "International Traffic in Arms Regulations: Modification of U.S. Munitions List Category XX ...",
34        "publication_date": "2026-09-18",
35        "type": "Rule",
36        "agencies": [
37          "State Department"
38        ],
39        "abstract": "The Department of State (the Department) amends the International Traffic in Arms Regulati ...",
40        "url": "https://www.federalregister.gov/documents/2026/09/18/2026-19211/international-traffi ..."
41      }
42    ],
43    "insights": {
44      "documents_analyzed": 25,
45      "document_mix": {
46        "Presidential Document": 1,
47        "Proposed Rule": 1,
48        "Rule": 2,
49        "Notice": 21
50      },
51      "activity_by_publication_day": {
52        "2026-09-18": 25
53      },
54      "top_agencies": [
55        {
56          "name": "Energy Department",
57          "documents": 7
58        },
59        {
60          "name": "Federal Energy Regulatory Commission",
61          "documents": 7
62        },
63        {
64          "name": "Federal Reserve System",
65          "documents": 2
66        },
67        {
68          "name": "Homeland Security Department",
69          "documents": 2
70        },
71        {
72          "name": "United States Sentencing Commission",
73          "documents": 2
74        }
75      ],
76      "abstract_coverage": "...",
77      "signals": {
78        "most_active_agency": "Energy Department",
79        "dominant_document_type": "Notice"
80      }
81    },
82    "provenance": "https://www.federalregister.gov/developers/documentation/api/v1",
83    "upgrade": {
84      "options": {
85        "pass": "$5 for 250 queries / 30 days",
86        "subscription": "$5/month for 2,000 queries"
87      },
88      "benefits": [
89        "up to 100 rows per query",
90        "date and document-type filters",
91        "aggregations and provenance",
92        "120 requests per minute"
93      ],
94      "access_start": "https://agentnative.cazimedia.com/v1/access/start"
95    },
96    "rate_limit": {
97      "limit": 20,
98      "remaining": 19,
99      "reset": 1789882740
100    }
101  }
102}

Its shape matters more than its content:

- data.rows holds three documents, each with a document number, title, date, type, agencies and a federalregister.gov URL.

- data.insights summarizes 25 documents: the mix of types and the top agencies. The 25 are counted, not returned.

- data.upgrade offers a $5 pass or subscription for up to 100 rows and date and type filters. The graph ignores it.

- data.rate_limit reports a limit of 20 calls. The graph ignores it too.

I first assumed the free tier returned the 25 newest documents, and it does not. Three rows is a small sample, and that shapes what this demo can honestly claim, which I come back to at the end.

The exchange, step by step

This sequence shows the remote agent being called, the documents coming back, and what the graph does with them next:

The remote agent appears only in the first two steps. Once its reply arrives, the graph works from the official Federal Register record, and the model never talks to the agent at all.

What the code is built with, and why

Yes, the graph is LangGraph. These are all the pieces the script imports or depends on, with the versions I ran:

PieceVersionWhat it does hereWhy this one
LangGraphlanggraph 1.2.11Defines the control flow as a graph: StateGraph, conditional edges, Send, a reducer, MemorySaver and interruptThe design needs parallel copies with a join, a conditional early stop and a pause that can resume, and LangGraph provides each of those
LangChain's Ollama integrationlangchain-ollama 1.1.0, with langchain-core 1.6.3ChatOllama talks to the local model, and with_structured_output returns a validated objectIt is the one small part of LangChain the graph needs. Nothing else from LangChain is imported
Pydantic2.13.5Defines the Score and Scores shapes the model must returnThe model's answer is checked against a schema before any code uses it
Ollama and qwen2.5:14bserver 0.34.2Runs the model on this machine and scores the documents and writes the noteNo API key, nothing leaves the machine, and the model was already installed. It returned valid structured output in every run I made
A2A Python SDKa2a-sdk 1.1.4Reads the Agent Card, works out which transport to use and sends the messageIt removes the hand-built JSON-RPC I used in part 2, which was there to show the wire format
httpx0.28.1Calls the federalregister.gov API for the verification stepAsync HTTP, and it is already a dependency of the A2A SDK

Not used: the prebuilt ReAct agent (create_react_agent, now create_agent in LangChain) and its tool-calling loop (the reasoning is under "Why the remote agent is called from a node" below), RemoteGraph and the LangGraph server, LangSmith, MCP, and the SDK's server side.

A2A does not require LangGraph. The protocol only defines what crosses the boundary between two agents, so the remote agent could be built on anything and my client could be a plain script. LangGraph is one way to write the client side.

Do you need LangGraph for this?

No. For three rows and one pause, the same steps fit in a plain async script: asyncio.gather for the verification, an if for the early stop and input() for the approval. That version would be shorter, and for a task this small it is arguably the better choice. LangGraph earns its place in three ways:

- The pause is a real stop. State is saved after every step, so interrupt ends the run and Command(resume=...) continues it, instead of a call blocked on the keyboard. In this script both halves run in one process, and MemorySaver keeps state in memory only, so I have shown the mechanism and not a run that survives a restart. That would need a persistent checkpointer.

- The graph can be drawn from the code. build_graph().get_graph() reported the nodes and edges, and I used them for the diagram below.

- Merging parallel results is declared, not written. The operator.add reducer says how the three verification results combine.

The cost is more concepts to learn and one more layer between you and the control flow. I judge the trade worth it here because this post is about orchestration, and LangGraph makes each part of it visible and named.

LangChain, LangGraph and Langflow

The three names look like one product and are not. LangChain and LangGraph are separate libraries from the same company, and one sits on top of the other. Langflow is a different kind of tool, mentioned here in passing because I use neither its canvas nor its runtime. None of the three is required to build an agent, as the last part of this section explains.

LangChainLangGraphLangflow
What it isBuilding blocks for LLM applications: model integrations, tools, prompts and ready-made agentsA framework and runtime for stateful agents, written as a graph of nodes joined by edgesA visual tool for building AI agents and workflows on a canvas, with a no-code or low-code interface
You work inCodeCodeA browser canvas
Who decides what runs nextMostly the agent loop: the model picks a tool, reads the result and repeatsYou, through the edges you draw, with agentic steps where you chooseThe flow you connect on the canvas
SuitsGetting an agent running quicklyPrecise control of every step, mixing fixed and model-driven steps, saved state and human approvalPrototyping multi-step or multi-agent applications quickly
In this postOnly ChatOllama and structured outputThe whole control flowNot used

The LangChain documentation puts the relationship plainly: "LangChain's agents are built on top of LangGraph. This allows us to take advantage of LangGraph's durable execution, human-in-the-loop support, persistence, and more." It describes create_agent as "a minimal, highly configurable agent harness", and describes LangGraph as "a low-level orchestration framework and runtime for building, managing, and deploying long-running, stateful agents." It also says LangGraph "can be used without LangChain". So the choice is not one library against the other. It is how much of the control flow you want a prebuilt harness to decide for you.

The pre-built agent has moved between them. When I first ran a LangGraph ReAct agent for this project, the import of create_react_agent printed a deprecation warning saying it "has been moved to langchain.agents" and to use create_agent instead. The idea is unchanged: a model that chooses tools in a loop. I wrote about that loop in my 2024 post on LangChain Agents, where an LLM decides which tools to call, runs them and repeats until the task is done. LangGraph draws the same loop as a cycle between an agent node and a tools node, which is the shape I contrast with this graph below.

I chose LangGraph directly for this script because the orchestration is the point. A prebuilt harness makes the "what next" decision inside a model loop, and I wanted those decisions in edges I can read, test and draw. I still use one LangChain piece, ChatOllama, because it is a convenient way to call a local model and get structured output back. The LangChain documentation recommends its higher-level agents for people getting started, and LangGraph for "precise control over every part of your agent's behavior."

Langflow sits in a different place. Where LangGraph is a graph you write in Python and can print as Mermaid text, Langflow lets you connect models, tools, prompts and memory as boxes on a canvas. I tried it in October 2024 in my post Langflow: I installed it with Python 3.10, started it with python -m langflow run, opened its canvas at http://127.0.0.1:7860 and imported a "Doc to Podcast" flow, which needed an openai_api_key variable. That is a good way to prototype quickly. I have not used Langflow for this task, and I have not checked how its current version works, so treat this paragraph as a pointer, not a comparison I have tested. If you want to see what a flow looks like before writing code, start there. If you need the flow to be tested, versioned and driven from code, as an orchestrator that must verify a third party's output does, a code-level graph like LangGraph is the closer fit.

You do not need a framework to build an agent

None of the three is required. An LLM-powered agent is a loop, and you can write that loop yourself. Send the conversation and a list of tool definitions to the model. If the reply asks for tool calls, run them, add the results to the conversation and go round again. If it does not, the reply is the answer. Put a cap on the number of trips so a confused model cannot loop forever.

That is the whole idea, and I built it without an agent framework such as LangChain or LangGraph in Claude Code, part 9, where I used Claude Code to create a local coding agent. Its tech stack is Next.js, TypeScript, Tailwind and Ollama running qwen2.5-coder:7b, and its core is one function, runAgent. The function follows the diagram above. It calls the model with the conversation and the tool definitions, runs any requested tools and feeds the results back, up to 20 iterations (MAX_TOOL_ITERATIONS). The agent has four tools: read_file, write_file, run_command and search_files. The function reports progress through a callback, so the same loop serves both the browser interface and the command line.

Writing the loop yourself also means owning its awkward cases. That post handles a model that does not support native function calling and puts its tool call as JSON inside the text of its reply, by scanning the text for it. Part 2 shows the other side of A2A working without a framework too: the remote helloworld agent has no agent framework and no model, and its whole "agent" is one function that echoes the request.

A framework is a convenience, not a requirement. What LangGraph adds to a loop like this is what this script uses it for: state saved between steps, a pause that can resume, parallel branches with a defined merge, and a graph you can draw from the code. If you do not need those, the loop above is enough. The question is whether the extras are worth another layer to learn, which is the same question I asked of this script under "Do you need LangGraph for this?" above.

What an orchestrator is responsible for

A2A describes the remote agent's side in detail: the card, the skills, the task states. The client that coordinates one or more agents has duties of its own, and most of them are easy to lose when the model is left to improvise.

Step 5 deserves emphasis. The remote agent is a third party, and the official samples' own README says to treat everything an external agent returns, including its card and messages, as untrusted input. The verification step in this graph is that advice made concrete.

PatternWhat it meansIn this post
SequentialStep B needs step A's outputDiscover, fetch, score, draft
Parallel fan-out and fan-inIndependent calls run together and a join waits for allOne verification per document
Conditional routingThe next step depends on a resultStop early if nothing is relevant
SupervisorOne agent routes work to several specialistsNot used
HandoffControl moves to another agent for the rest of the conversationNot used
Human in the loopThe run pauses for approval before an irreversible stepApprove the note before it can be sent

The graph

The control flow is a LangGraph StateGraph. This diagram shows the nodes, the edge types, the shared state, the checkpointer and the pause.

PartWhat it isWhere it appears here
NodeA function that receives the State and returns updates to itThe six named boxes
Fixed edgeAlways go from A to BThe thick arrows
Conditional edgeA function chooses the destination at run time, and can return severalfan_out and route
SendStart a node with its own private input, one copy per itemOne verify_document per row
ReducerA rule for merging updates to a field from parallel nodesoperator.add on checked
CheckpointerSaves the State after each step so a run can stop and continueMemorySaver, which makes the pause possible
InterruptA node stops the run and returns a value to the callerhuman_review
CycleAn edge back to an earlier nodeNone in this graph

A ReAct agent, the pattern behind LangGraph's older create_react_agent (now moved to LangChain as create_agent), is the opposite shape: agent and tools are joined in a cycle and the model decides how many times around it goes. This graph has no cycle. Every node runs once, in an order I fixed. A cycle is one more edge if a loop is what you want, for example retrying a failed call or waiting while a remote task is still working, and nothing in this run needed one.

Why the remote agent is called from a node

fetch_briefing is an ordinary async function that calls the A2A client. The model never sees the remote agent, so it cannot skip the call, run it early or fill in its arguments. I chose this because fetching the briefing is always required and nothing about it needs judgment. The alternative is to expose the agent as a tool, which suits a supervisor that must pick among several agents, and the next subsection compares the two. I tried that pattern in an earlier draft of this post with a different agent and a small local model, and the model skipped a tool, called one before its inputs existed and passed wrong values. That is evidence from one model, and a stronger one might behave better, but a node removes the question.

The model has two jobs here, scoring relevance and writing a paragraph. It never sees the remote agent's text. It sees fields taken from the official federalregister.gov record, and code, not the model, adds the links to the note.

Remote agents and sub-agents as tools

Wrapping an agent as a tool is a documented and common pattern, and it is the main alternative to what this script does. LangChain's multi-agent guide lists it first, under the name subagents: "A main agent coordinates subagents as tools. All routing passes through the main agent, which decides when and how to invoke each subagent." It says the pattern suits parallel work and large contexts, because each sub-agent works "in isolation with only its relevant context." The guide does not mention remote agents or A2A, so it does not say whether a sub-agent runs in the same process or across a network.

In part 1 I measured this pattern against a router and against one big prompt for a different job, looking things up in six documents I had written. Subagents were no more accurate there than a router or a skills agent, and they used the most prompt tokens of the three designs that fit the model's window. That result is about documents you can load into one prompt. A remote agent is the case the pattern is really for, because you cannot load its data or its logic into your own context at all, and the boundary is real.

A sub-agent is an agent with its own loop and its own context, called by a main agent. From the main agent's side it is a name, a description and some arguments, and what comes back is a result. That shape is the same whether the sub-agent runs locally or is a remote A2A agent behind a wrapper. The network only adds what this post has been dealing with: finding the agent, deciding whether to trust its reply, and handling task states and rate limits.

Graph node (this script)Sub-agent as a toolHandoff
Who decides to call itThe graph, through its edgesThe main agent's modelThe current agent passes control, through a tool call
What the model seesNothing about the remote agentA name, a description and arguments, then the resultThe conversation moves to the other agent
SuitsA step that is always needed, checked before use, or run in parallelA main agent that must choose among several agents, calling well-defined skillsA conversation the other agent should carry on
Main riskThe graph cannot adapt on its own to a case you did not drawThe model skips it, misorders it or fills in wrong argumentsThe original agent loses control of the outcome

The handoff column and the "main risk" row are my own summary, not quotations from the guide. The A2A documentation accepts tool-style wrapping and marks its limit. On its page about A2A and MCP, exposing an agent as a tool "works best when the skills are well-defined and can be called in a tool-like, stateless way", and "A2A's main strength is its support for flexible, stateful, collaborative interactions that go beyond a typical tool call." The two agents in this series are the tool-like kind: one well-defined request and one answer. Wrapping either as a tool would work.

Three points are my own inferences, not documented claims:

- A wrapper hides the parts of A2A beyond a tool call. A tool returns one result. Streaming, multi-turn exchanges and task states such as input-required have to be squeezed into that result, or into an error the model can read. fetch_briefing raises on failed, rejected, input-required and auth-required states, and a tool wrapper would have to decide what the model should be told instead.

- A wrapper decides what the model reads. If the remote agent's text goes straight back to the model, an unverified third party is writing into the model's context. This script filters the reply and verifies each row first, and a tool wrapper should do the same before returning anything.

- The patterns combine. The same guide says "You can mix patterns! For example, a subagents architecture can invoke tools that invoke custom workflows." A graph like this one could be a single tool of a larger supervisor.

My rule of thumb is to use a node when the call is always needed and nothing about it needs judgment, and a tool when the model genuinely has to choose which agent to call. My own experience is one small local model over a few runs, in which it skipped a tool and misordered calls. That is not a measurement across models, and a stronger model may need the safeguards less.

The code

regulatory_watch.py:

PYTHON
1import asyncio
2import json
3import operator
4import re
5import sys
6import uuid
7from typing import Annotated, TypedDict
8
9import httpx
10from a2a.client import A2ACardResolver, ClientConfig, create_client
11from a2a.types import Message, Part, Role, SendMessageRequest, TaskState
12from langchain_ollama import ChatOllama
13from langgraph.checkpoint.memory import MemorySaver
14from langgraph.graph import END, START, StateGraph
15from langgraph.types import Command, Send, interrupt
16from pydantic import BaseModel
17
18AGENT_URL = "https://agentnative.cazimedia.com"
19SKILL_ID = "federal-register-query"
20BRIEFING_PROMPT = "Give me a free current Federal Register briefing with provenance."
21OFFICIAL_API = "https://www.federalregister.gov/api/v1/documents/{number}.json"
22OFFICIAL_FIELDS = ["document_number", "title", "type", "publication_date", "html_url", "agencies", "abstract"]
23MODEL = "qwen2.5:14b"
24DEFAULT_ROLE = "export compliance officer at a manufacturer of underwater and defense electronics"
25RECIPIENT = "Dana Okafor, general counsel"
26
27BAD_STATES = {TaskState.TASK_STATE_FAILED, TaskState.TASK_STATE_REJECTED,
28              TaskState.TASK_STATE_INPUT_REQUIRED, TaskState.TASK_STATE_AUTH_REQUIRED}
29
30
31class State(TypedDict):
32    role: str
33    card: dict
34    rows: list[dict]
35    insights: dict
36    checked: Annotated[list[dict], operator.add]  # parallel branches append here
37    picks: list[dict]
38    note: str
39    decision: str
40
41
42class Score(BaseModel):
43    document_number: str
44    relevance: int
45    reason: str
46
47
48class Scores(BaseModel):
49    scores: list[Score]
50
51
52def reply_text(event) -> str:
53    kind = event.WhichOneof("payload")
54    if kind == "message":
55        return "".join(p.text for p in event.message.parts)
56    if kind == "task":
57        if event.task.status.state in BAD_STATES:
58            raise RuntimeError(f"task ended {TaskState.Name(event.task.status.state)}")
59        return "".join(p.text for a in event.task.artifacts for p in a.parts)
60    return ""
61
62
63async def discover(state: State) -> dict:
64    async with httpx.AsyncClient() as http:
65        card = await A2ACardResolver(http, AGENT_URL).get_agent_card()
66    skills = [s.id for s in card.skills]
67    if SKILL_ID not in skills:
68        raise RuntimeError(f"{card.name} does not advertise {SKILL_ID}")
69    iface = card.supported_interfaces[0]
70    print(f"  [discover] {card.name} v{card.version}: {iface.protocol_binding} "
71          f"{iface.protocol_version}, {len(skills)} skills, {SKILL_ID} found")
72    return {"card": {"name": card.name, "skills": skills}}
73
74
75async def fetch_briefing(state: State) -> dict:
76    client = await create_client(AGENT_URL, ClientConfig(streaming=False))
77    text = ""
78    try:
79        message = Message(message_id=str(uuid.uuid4()), role=Role.ROLE_USER,
80                          parts=[Part(text=BRIEFING_PROMPT)])
81        async for event in client.send_message(SendMessageRequest(message=message)):
82            text += reply_text(event)
83    finally:
84        await client.close()
85
86    reply = json.loads(text)
87    if reply.get("type") != "federal-register-briefing" or reply.get("status") != 200:
88        raise RuntimeError(f"unexpected reply: {text[:200]}")
89    data = reply["data"]
90
91    # The agent is an unverified third party: keep only the fields this graph needs,
92    # in the shape it expects. The upgrade offers and everything else are dropped.
93    rows = []
94    for r in data["rows"]:
95        if (re.fullmatch(r"\d{4}-\d{4,6}", str(r.get("document_number", "")))
96                and str(r.get("url", "")).startswith("https://www.federalregister.gov/")):
97            rows.append({"document_number": r["document_number"], "title": str(r.get("title", ""))[:300],
98                         "type": str(r.get("type", "")), "publication_date": str(r.get("publication_date", "")),
99                         "url": r["url"]})
100    ins = data.get("insights", {})
101    insights = {"documents_analyzed": ins.get("documents_analyzed"), "document_mix": ins.get("document_mix"),
102                "top_agencies": ins.get("top_agencies", [])[:5]}
103    remaining = data.get("rate_limit", {}).get("remaining")
104    print(f"  [fetch_briefing] {len(rows)} rows kept of {len(data['rows'])}; "
105          f"insights cover {insights['documents_analyzed']} documents; free-tier calls remaining: {remaining}")
106    return {"rows": rows, "insights": insights}
107
108
109def fan_out(state: State) -> list[Send]:
110    return [Send("verify_document", {"row": r}) for r in state["rows"]]
111
112
113def squash(s: str) -> str:
114    return re.sub(r"\s+", " ", s).strip()
115
116
117async def verify_document(payload: dict) -> dict:
118    row = payload["row"]
119    params = [("fields[]", f) for f in OFFICIAL_FIELDS]
120    async with httpx.AsyncClient(timeout=30) as http:
121        resp = await http.get(OFFICIAL_API.format(number=row["document_number"]), params=params)
122    problems, official = [], None
123    if resp.status_code == 404:
124        problems.append("not found at federalregister.gov")
125    else:
126        resp.raise_for_status()
127        official = resp.json()
128        for field, official_field in (("title", "title"), ("type", "type"),
129                                      ("publication_date", "publication_date"), ("url", "html_url")):
130            if squash(row[field]) != squash(official[official_field]):
131                problems.append(f"{field} differs from the official record")
132    ok = not problems
133    print(f"  [verify_document] {row['document_number']}: {'matches the official record' if ok else '; '.join(problems)}")
134    result = {"document_number": row["document_number"], "ok": ok, "problems": problems}
135    if ok:
136        result["official"] = {
137            "title": official["title"], "type": official["type"], "date": official["publication_date"],
138            "agencies": [a["name"] for a in official["agencies"]], "url": official["html_url"],
139            "abstract": (official.get("abstract") or "")[:700]}
140    return {"checked": [result]}
141
142
143def score_relevance(state: State) -> dict:
144    good = {c["document_number"]: c["official"] for c in state["checked"] if c["ok"]}
145    if not good:
146        return {"picks": []}
147    docs = "\n\n".join(
148        f"{n}: {d['title']}\nType: {d['type']}. Agencies: {', '.join(d['agencies'])}. Published {d['date']}.\n"
149        f"Abstract: {d['abstract'] or '(none published)'}" for n, d in good.items())
150    model = ChatOllama(model=MODEL, temperature=0).with_structured_output(Scores)
151    result = model.invoke(
152        f"I am a {state['role']}. Below are new Federal Register documents. For each one, give a "
153        "relevance score from 0 (irrelevant to my work) to 10 (I must read it today) and a one-sentence "
154        "reason that refers to my role. Use each document_number exactly as given.\n\n" + docs)
155    picks = []
156    for s in result.scores:
157        if s.document_number in good and s.relevance >= 6:
158            picks.append({**good[s.document_number], "document_number": s.document_number,
159                          "relevance": min(s.relevance, 10), "reason": s.reason})
160    for s in result.scores:
161        print(f"  [score_relevance] {s.document_number}: {s.relevance}/10 - {s.reason}")
162    return {"picks": sorted(picks, key=lambda p: -p["relevance"])}
163
164
165def route(state: State) -> str:
166    return "draft_note" if state["picks"] else "nothing_relevant"
167
168
169def draft_note(state: State) -> dict:
170    items = "\n".join(f"- {p['title']} ({p['type']}, {p['date']}). Why it matters: {p['reason']}"
171                      for p in state["picks"])
172    body = ChatOllama(model=MODEL, temperature=0).invoke(
173        f"Write the body of a short note (under 120 words) from Neil, a {state['role']}, to {RECIPIENT}, "
174        "flagging the Federal Register documents below. Do not include links, document numbers, a greeting "
175        f"or a sign-off, and do not add facts that are not listed.\n\n{items}").content.strip()
176    # Links come from the verified official records, never from the model.
177    sources = "\n".join(f"- {p['title']} ({p['type']}, {p['date']}): {p['url']}" for p in state["picks"])
178    note = f"Dear {RECIPIENT.split(',')[0].split()[0]},\n\n{body}\n\nSources:\n{sources}\n\nBest regards,\nNeil"
179    print("  [draft_note] drafted")
180    return {"note": note}
181
182
183def human_review(state: State) -> dict:
184    return {"decision": interrupt({"note": state["note"], "question": "Send this note? (approve / reject)"})}
185
186
187def build_graph():
188    g = StateGraph(State)
189    for name, fn in (("discover", discover), ("fetch_briefing", fetch_briefing),
190                     ("verify_document", verify_document), ("score_relevance", score_relevance),
191                     ("draft_note", draft_note), ("human_review", human_review)):
192        g.add_node(name, fn)
193    g.add_edge(START, "discover")
194    g.add_edge("discover", "fetch_briefing")
195    g.add_conditional_edges("fetch_briefing", fan_out, ["verify_document"])
196    g.add_edge("verify_document", "score_relevance")
197    g.add_conditional_edges("score_relevance", route, {"draft_note": "draft_note", "nothing_relevant": END})
198    g.add_edge("draft_note", "human_review")
199    g.add_edge("human_review", END)
200    return g.compile(checkpointer=MemorySaver())
201
202
203async def run_until_pause(graph, payload, config):
204    async for update in graph.astream(payload, config, stream_mode="updates"):
205        if "__interrupt__" in update:
206            return update["__interrupt__"][0].value
207    return None
208
209
210async def tamper_test() -> None:
211    real = {"document_number": "2026-19211", "title": "International Traffic in Arms Regulations: "
212            "Modification of U.S. Munitions List Category XX(a)", "type": "Rule",
213            "publication_date": "2026-09-18", "url": "https://www.federalregister.gov/documents/2026/09/18/"
214            "2026-19211/international-traffic-in-arms-regulations-modification-of-us-munitions-list-category-xxa"}
215    for label, row in (("untouched row", real),
216                       ("title altered", {**real, "title": "Repeal of all export controls"}),
217                       ("invented document", {**real, "document_number": "2026-99999"})):
218        print(f"{label}:")
219        await verify_document({"row": row})
220
221
222async def main(role: str) -> None:
223    graph = build_graph()
224    config = {"configurable": {"thread_id": "watch-1"}}
225    print(f"role: {role}\n--- run ---")
226    pending = await run_until_pause(graph, {"role": role, "checked": []}, config)
227    if pending is None:
228        print("\n--- finished without pausing: no document was relevant, nothing drafted ---")
229        return
230    print("\n--- paused for human review ---")
231    print(pending["note"])
232    answer = input(f"\n{pending['question']} ").strip() or "reject"
233    print("\n--- resumed ---")
234    await run_until_pause(graph, Command(resume=answer), config)
235    final = (await graph.aget_state(config)).values
236    print(f"decision recorded: {final['decision']}")
237
238
239if __name__ == "__main__":
240    if "--tamper-test" in sys.argv:
241        asyncio.run(tamper_test())
242    elif "--graph" in sys.argv:
243        print(build_graph().get_graph().draw_mermaid())
244    else:
245        asyncio.run(main(next((a for a in sys.argv[1:] if not a.startswith("--")), DEFAULT_ROLE)))

Running it with uv

uv is a fast Python package and version manager. It can download a Python interpreter of its own, which matters here: the script needs Python 3.10 or newer, and the system Python on my Mac was 3.9.6, as I found in the first post. On macOS I installed it with Homebrew:

BASH
1brew install uv
Before you start

- Ollama must be running with the model pulled. The scoring step is the first place the script uses the model, so a missing model or a stopped server would show up there:

BASH
1  ollama pull qwen2.5:14b
2  curl http://localhost:11434/api/version

The curl prints Ollama's version if the server is up. The model is about 9 GB.

- Make a folder and save two files in it. Save the code above as regulatory_watch.py, and save this as requirements.txt:

CODE
1  langgraph
2  langchain-ollama
3  httpx
4  a2a-sdk

These are unpinned. The versions I ran are in the table above, and you can pin them, for example langgraph==1.2.11, if you want the same behavior later.

Option 1: a virtual environment
BASH
1uv venv --python 3.12 .venv
2source .venv/bin/activate
3uv pip install -r requirements.txt

uv venv --python 3.12 .venv creates an isolated environment in a folder called .venv, and downloads a CPython 3.12 build if the machine does not already have one. source .venv/bin/activate makes that environment the current one for this terminal window. uv pip install -r requirements.txt installs the four packages, and everything they depend on, into it. It uses uv's pip-compatible interface, so it reads an ordinary requirements.txt.

Then run the script. Start with the check that needs no model and no call to the remote agent, only the internet:

BASH
1python regulatory_watch.py --tamper-test

Then the graph itself:

BASH
1python regulatory_watch.py                                      # the default role
2python regulatory_watch.py "marine biologist restoring coral reefs"
3echo approve | python regulatory_watch.py                       # answers the approval prompt for you
4python regulatory_watch.py --graph                              # prints the compiled graph as Mermaid text

The role is the first argument that does not start with --, so quote it if it has spaces. At the pause the script prints the draft note and waits: type approve or reject and press Return. The echo approve | form feeds that answer in, which is how I captured the run below. If nothing is relevant to the role, the run ends before the pause, as in my second run.

Option 2: uv run, with nothing to activate
BASH
1uv run --python 3.12 --with-requirements requirements.txt regulatory_watch.py --tamper-test

uv run builds a temporary environment from the requirements file, keeps it in uv's cache, runs the script inside it and leaves nothing installed on the rest of the machine. Everything after the script name is passed to the script, so the same arguments work:

BASH
1uv run --python 3.12 --with-requirements requirements.txt regulatory_watch.py "marine biologist restoring coral reefs"

I ran the --tamper-test command both ways in a clean folder and it produced the output shown below each time. uv also has a project workflow, uv init and uv add, which records dependencies in a pyproject.toml. I did not use it here.

What to expect

- The output will differ from mine. The briefing changes every day, and a model at temperature 0 is repeatable for a given input, not across different inputs.

- The remote agent is rate limited. Its reply reported a limit of 20 calls. If it stops returning a briefing, the script stops with an error in fetch_briefing instead of carrying on with bad data.

- The first run is slower. uv downloads Python and the packages, and Ollama loads the model into memory.

Runs

Before running the whole graph, I tested the verification step alone, since a check that has never failed proves little. I fed it one untouched row, one with an altered title and one invented document, using no model and no calls to the remote agent:

CODE
1untouched row:
2  [verify_document] 2026-19211: matches the official record
3title altered:
4  [verify_document] 2026-19211: title differs from the official record
5invented document:
6  [verify_document] 2026-99999: not found at federalregister.gov

Then the whole graph, with the role of an export compliance officer, typing approve at the pause:

CODE
1role: export compliance officer at a manufacturer of underwater and defense electronics
2--- run ---
3  [discover] AgentNative Data Exchange v0.3.0: JSONRPC 1.0, 6 skills, federal-register-query found
4  [fetch_briefing] 3 rows kept of 3; insights cover 25 documents; free-tier calls remaining: 19
5  [verify_document] 2026-19222: matches the official record
6  [verify_document] 2026-19211: matches the official record
7  [verify_document] 2026-19251: matches the official record
8  [score_relevance] 2026-19251: 0/10 - This document pertains to drug transit and production and is not directly related to underwater or defense electronics.
9  [score_relevance] 2026-19222: 0/10 - This document concerns employment regulations and does not pertain to export compliance for underwater or defense electronics.
10  [score_relevance] 2026-19211: 10/10 - This document directly affects the export regulations for uncrewed underwater vehicles, which is highly relevant to my role in export compliance for defense electronics.
11  [draft_note] drafted
12
13--- paused for human review ---
14Dear Dana,
15
16Please be advised of a recent publication in the Federal Register regarding the International Traffic in Arms Regulations: Modification of U.S. Munitions List Category XX(a). This modification is crucial as it directly impacts the export regulations for uncrewed underwater vehicles, an area highly relevant to our export compliance for defense electronics. I recommend we review the implications of this change on our current export practices and ensure compliance with the updated regulations.
17
18Sources:
19- International Traffic in Arms Regulations: Modification of U.S. Munitions List Category XX(a) (Rule, 2026-09-18): https://www.federalregister.gov/documents/2026/09/18/2026-19211/international-traffic-in-arms-regulations-modification-of-us-munitions-list-category-xxa
20
21Best regards,
22Neil
23
24Send this note? (approve / reject) approve
25--- resumed ---
26decision recorded: approve

The three verify_document lines run in parallel and finish in whatever order they finish: the order differed between my two full runs. All three rows matched the official record, so the model saw all three. It scored the two irrelevant documents 0 and the International Traffic in Arms Regulations rule 10, which fits the role, since the rule removes certain uncrewed underwater vehicles from the U.S. Munitions List. The note's links come from the verified record.

Then the same graph with a role that should match nothing:

CODE
1role: marine biologist restoring coral reefs
2--- run ---
3  [discover] AgentNative Data Exchange v0.3.0: JSONRPC 1.0, 6 skills, federal-register-query found
4  [fetch_briefing] 3 rows kept of 3; insights cover 25 documents; free-tier calls remaining: 19
5  [verify_document] 2026-19211: matches the official record
6  [verify_document] 2026-19251: matches the official record
7  [verify_document] 2026-19222: matches the official record
8  [score_relevance] 2026-19251: 0/10 - This document is about drug transit and production and has no direct relation to marine biology or coral reef restoration.
9  [score_relevance] 2026-19222: 0/10 - This document pertains to employment regulations and does not impact marine biology or coral reef restoration efforts.
10  [score_relevance] 2026-19211: 2/10 - This document modifies regulations on uncrewed underwater vehicles, which could indirectly affect marine research and conservation efforts, but is not directly relevant to coral reef restoration.
11
12--- finished without pausing: no document was relevant, nothing drafted ---

This time the conditional edge sent the run to __end__ before drafting, so nothing was written and there was nothing to approve. The model gave the ITAR rule 2 out of 10 for a marine biologist, below the threshold of 6 that I set in code.

What this does and does not show

- The agent adds little for this data. The Federal Register publishes a free, keyless API, which is what the verification step calls. The three free rows are ones I could have fetched from it directly. The agent's value is in its paid tier of filters, up to 100 rows and aggregations, which I did not use. What the demo does show is the orchestration pattern: discover, delegate, verify, merge, hand back to a person.

- Reply handling was only partly exercised. The agent answered with a plain message, so the task-state branch (failed, rejected, input-required, auth-required) never ran. The agent also advertises a durable task for making datasets query-ready, which I did not try.

- No contextId was sent. Nothing here needs conversation state. The LangChain A2A documentation describes how an Agent Server maps a contextId to a LangGraph thread_id, which matters for agents that remember earlier turns. The thread_id in my checkpointer is a separate, local identifier.

- RemoteGraph was not used. It calls a graph deployed on a LangGraph server over that server's own API, and this agent is not one.

- The scoring rests on three documents. A model choosing among three rows is a thin test, and its scores are one small local model's opinion. I set the threshold of 6 by hand.

- The listing is unclaimed. Its owner has not verified it with the registry, so I treated its replies as untrusted and ignored the upgrade offers.

- The free tier is limited. Its own rate_limit field reported a limit of 20 calls, and I made a handful of calls in total.