Agent Orchestration (Part 3)
Neil Haddley • September 24, 2026
Finding a real agent through a2a-registry.org, then orchestrating it from a LangGraph graph with a local Ollama model: discover the Agent Card, delegate over A2A, verify the reply against the official source, and pause for human approval
This is part 3 of a series on agent orchestration. Part 1 measures four LangChain patterns for agents you build yourself, and part 2 introduces the Agent2Agent protocol.
My previous post, part 2, ran a "hello world" agent I wrote myself, to see the protocol's mechanics with nothing else in the way. This post uses an agent I did not write. I found it in a public registry, a2a-registry.org. In the LangGraph graph I fetch its Agent Card again on every run, so the run checks what the agent advertises today and not what the registry listed when I looked, and then I call it. The graph runs on my own machine with a model hosted by Ollama. The graph is the point: it shows what an orchestrator is responsible for when it depends on an agent it does not control.
The task is a regulatory watch. The graph asks the remote agent for a briefing of new Federal Register documents, checks every row against the official source, has a local model score each verified document for a stated job role, drafts a short note about the relevant ones, and stops for my approval before anything would be sent.
Every output below is from a real run against live services on 20 September 2026. The briefing changes daily, so a run today will differ.
a2a-registry.org
An Agent Card only helps once you know where to fetch it. a2a-registry.org covers the step before that, discovering agents you did not already know about. It describes itself as "the definitive directory for the Agentic Web" and, at the time of writing, lists 301 agents, 74 of them verified. Verification means the owner proved control through a DNS TXT record, a linked GitHub account or a GoDaddy ANS name. Adding a listing needs no account: you paste an agent's URL and the registry reads the Agent Card.
The registry is an A2A agent itself. Its own card declares a JSON-RPC interface with a search_agents skill that takes a natural-language query. It also declares bearer-token authentication, and an unauthenticated call returns HTTP 401 with a pointer to a token sign-up page. I did not create an account, so I browsed the categories by hand instead.
Listings vary a great deal, and the registry's own health status can be stale. Four agents I looked at all showed the identical "14 consecutive health check failures", and a direct request for one of their cards succeeded, which suggests the registry's checker had stalled, not that four agents failed together. Before choosing, I checked each candidate's card, or its registry page where I did not fetch the card, for two things: which transport it declares, and whether it charges.
| Agent | How I checked | Result |
|---|---|---|
| AgentNative Data Exchange | Card fetched, several calls made | JSON-RPC 1.0, declares no auth, free sample tier. Chosen |
| CharitySense | Card fetched | Transport is OPENAPI: a REST API described with an A2A-shaped card, not an A2A message exchange |
| ForgeMesh travel and fares | Registry page | Every call is paid in USDC through x402 crypto payments |
| The registry's own agent | Card fetched, call returned 401 | Needs an account token |
The agent: AgentNative Data Exchange
The registry listing points at an Agent Card, and the first thing the graph does is fetch it.
The Agent Card
This is the card exactly as the agent serves it at https://agentnative.cazimedia.com/.well-known/agent-card.json. I fetched it with curl and pretty-printed it:
JSON
1{ 2 "name": "AgentNative Data Exchange", 3 "description": "Normalized official government and public datasets across federal, state and city agencies, census, police, education and schools, with provenance, aggregations, insights and free samples.", 4 "url": "https://agentnative.cazimedia.com/a2a", 5 "version": "0.3.0", 6 "supportedInterfaces": [ 7 { 8 "url": "https://agentnative.cazimedia.com/a2a", 9 "protocolBinding": "JSONRPC", 10 "protocolVersion": "1.0" 11 }, 12 { 13 "url": "https://agentnative.cazimedia.com", 14 "protocolBinding": "HTTP+JSON", 15 "protocolVersion": "1.0" 16 } 17 ], 18 "capabilities": { 19 "streaming": false, 20 "pushNotifications": false, 21 "extendedAgentCard": false 22 }, 23 "defaultInputModes": [ 24 "text/plain", 25 "application/json" 26 ], 27 "defaultOutputModes": [ 28 "text/plain", 29 "application/json" 30 ], 31 "skills": [ 32 { 33 "id": "public-data-search", 34 "name": "Search official government and public datasets", 35 "description": "Discover federal, state, city, census, police, education and school data with official provenance.", 36 "tags": [ 37 "government data", 38 "public datasets", 39 "statistics", 40 "provenance" 41 ], 42 "examples": [ 43 "Find official datasets about public schools." 44 ] 45 }, 46 { 47 "id": "on-demand-materialization", 48 "name": "Make discovered data query-ready", 49 "description": "Prioritize a compatible official dataset for detached ingestion and poll the durable task through completion.", 50 "tags": [ 51 "materialization", 52 "async task", 53 "data import" 54 ], 55 "examples": [ 56 "Materialize disc_0123456789abcdef01234567 and tell me when it is query-ready." 57 ] 58 }, 59 { 60 "id": "imported-data-sample", 61 "name": "Sample normalized government datasets free", 62 "description": "Inspect official rows, summaries, deterministic insights, freshness and provenance before payment.", 63 "tags": [ 64 "free sample", 65 "open data", 66 "normalized records" 67 ], 68 "examples": [ 69 "Show me useful free data with provenance." 70 ] 71 }, 72 { 73 "id": "imported-data-query", 74 "name": "Query and aggregate official datasets", 75 "description": "Filter, group and aggregate normalized public records across government agencies.", 76 "tags": [ 77 "aggregation", 78 "analysis", 79 "official records" 80 ], 81 "examples": [ 82 "Group these public records by agency." 83 ] 84 }, 85 { 86 "id": "coverage-status", 87 "name": "Inspect source coverage", 88 "description": "Inspect current catalog traversal, materialization, queue, failure, and publication freshness state.", 89 "tags": [ 90 "coverage", 91 "freshness", 92 "data sources" 93 ], 94 "examples": [ 95 "Which official sources are currently queryable?" 96 ] 97 }, 98 { 99 "id": "federal-register-query", 100 "name": "Query Federal Register", 101 "description": "Query normalized rules, proposed rules, notices, and presidential documents with bounded date and type filters.", 102 "tags": [ 103 "Federal Register", 104 "regulations", 105 "rules", 106 "notices" 107 ], 108 "examples": [ 109 "Give me a free current Federal Register briefing with provenance." 110 ] 111 } 112 ] 113}
What matters in it:
- supportedInterfaces lists two ways to reach the agent: JSONRPC at https://agentnative.cazimedia.com/a2a and HTTP+JSON at the base URL, both at protocol version 1.0. The graph's discover step prints the first entry, JSONRPC 1.0, and the hand-made call below uses the JSON-RPC one.
- capabilities.streaming is false, so a call returns one message and not a stream of updates.
- There is no securitySchemes key, so the card declares no authentication.
- Six skills are listed. The one this post uses is federal-register-query, and discover stops the run if that id is missing. The card gives an example prompt for it and no input schema, so the example sentence is the only guidance on how to ask. That is why the graph sends that exact sentence. A different free-text prompt I tried, a search for datasets about public schools, returned a capabilities menu and no data.
- on-demand-materialization advertises a "durable task" you poll to completion, which would exercise A2A's task states. I did not use it.
Calling it
The a2a-sdk client is given nothing beyond the base URL:
PYTHON
1client = await create_client("https://agentnative.cazimedia.com", ClientConfig(streaming=False)) 2message = Message(message_id=str(uuid.uuid4()), role=Role.ROLE_USER, 3 parts=[Part(text="Give me a free current Federal Register briefing with provenance.")]) 4async for event in client.send_message(SendMessageRequest(message=message)): 5 ...
To see the wire format, I made the same call by hand once with curl, using the JSON-RPC interface from the card and the A2A-Version: 1.0 header:
BASH
1curl -X POST https://agentnative.cazimedia.com/a2a \ 2 -H "Content-Type: application/json" -H "A2A-Version: 1.0" \ 3 -d '{"jsonrpc":"2.0","id":1,"method":"SendMessage","params":{"message":{"role":"ROLE_USER","message_id":"wire-check-1","parts":[{"text":"Give me a free current Federal Register briefing with provenance."}]}}}'
The response is a JSON-RPC envelope holding one A2A message, from the ROLE_AGENT, with a single text part:
JSON
1{"jsonrpc": "2.0", "id": 1, "result": {"message": {"messageId": "...", "role": "ROLE_AGENT", "parts": [{"text": "<a JSON document, shown next>"}]}}}
The reply
The text part is itself a JSON document. Here it is, decoded and abridged: I shortened the long title, abstract and URL strings and collapsed one field, and changed nothing else.
JSON
1{ 2 "type": "federal-register-briefing", 3 "status": 200, 4 "data": { 5 "dataset": "us-federal-register-documents", 6 "sample": true, 7 "sample_basis": "Aggregates cover the 25 newest documents; three example records are included.", 8 "rows": [ 9 { 10 "document_number": "2026-19251", 11 "title": "Presidential Determination on Major Drug Transit or Major Illicit Drug Producing Countries ...", 12 "publication_date": "2026-09-18", 13 "type": "Presidential Document", 14 "agencies": [ 15 "Executive Office of the President" 16 ], 17 "abstract": null, 18 "url": "https://www.federalregister.gov/documents/2026/09/18/2026-19251/presidential-determi ..." 19 }, 20 { 21 "document_number": "2026-19222", 22 "title": "Employment in the Excepted Service", 23 "publication_date": "2026-09-18", 24 "type": "Proposed Rule", 25 "agencies": [ 26 "Personnel Management Office" 27 ], 28 "abstract": "The Office of Personnel Management (OPM) proposes to amend its regulations governing the e ...", 29 "url": "https://www.federalregister.gov/documents/2026/09/18/2026-19222/employment-in-the-ex ..." 30 }, 31 { 32 "document_number": "2026-19211", 33 "title": "International Traffic in Arms Regulations: Modification of U.S. Munitions List Category XX ...", 34 "publication_date": "2026-09-18", 35 "type": "Rule", 36 "agencies": [ 37 "State Department" 38 ], 39 "abstract": "The Department of State (the Department) amends the International Traffic in Arms Regulati ...", 40 "url": "https://www.federalregister.gov/documents/2026/09/18/2026-19211/international-traffi ..." 41 } 42 ], 43 "insights": { 44 "documents_analyzed": 25, 45 "document_mix": { 46 "Presidential Document": 1, 47 "Proposed Rule": 1, 48 "Rule": 2, 49 "Notice": 21 50 }, 51 "activity_by_publication_day": { 52 "2026-09-18": 25 53 }, 54 "top_agencies": [ 55 { 56 "name": "Energy Department", 57 "documents": 7 58 }, 59 { 60 "name": "Federal Energy Regulatory Commission", 61 "documents": 7 62 }, 63 { 64 "name": "Federal Reserve System", 65 "documents": 2 66 }, 67 { 68 "name": "Homeland Security Department", 69 "documents": 2 70 }, 71 { 72 "name": "United States Sentencing Commission", 73 "documents": 2 74 } 75 ], 76 "abstract_coverage": "...", 77 "signals": { 78 "most_active_agency": "Energy Department", 79 "dominant_document_type": "Notice" 80 } 81 }, 82 "provenance": "https://www.federalregister.gov/developers/documentation/api/v1", 83 "upgrade": { 84 "options": { 85 "pass": "$5 for 250 queries / 30 days", 86 "subscription": "$5/month for 2,000 queries" 87 }, 88 "benefits": [ 89 "up to 100 rows per query", 90 "date and document-type filters", 91 "aggregations and provenance", 92 "120 requests per minute" 93 ], 94 "access_start": "https://agentnative.cazimedia.com/v1/access/start" 95 }, 96 "rate_limit": { 97 "limit": 20, 98 "remaining": 19, 99 "reset": 1789882740 100 } 101 } 102}
Its shape matters more than its content:
- data.rows holds three documents, each with a document number, title, date, type, agencies and a federalregister.gov URL.
- data.insights summarizes 25 documents: the mix of types and the top agencies. The 25 are counted, not returned.
- data.upgrade offers a $5 pass or subscription for up to 100 rows and date and type filters. The graph ignores it.
- data.rate_limit reports a limit of 20 calls. The graph ignores it too.
I first assumed the free tier returned the 25 newest documents, and it does not. Three rows is a small sample, and that shapes what this demo can honestly claim, which I come back to at the end.
The exchange, step by step
This sequence shows the remote agent being called, the documents coming back, and what the graph does with them next:
The remote agent appears only in the first two steps. Once its reply arrives, the graph works from the official Federal Register record, and the model never talks to the agent at all.
What the code is built with, and why
Yes, the graph is LangGraph. These are all the pieces the script imports or depends on, with the versions I ran:
| Piece | Version | What it does here | Why this one |
|---|---|---|---|
| LangGraph | langgraph 1.2.11 | Defines the control flow as a graph: StateGraph, conditional edges, Send, a reducer, MemorySaver and interrupt | The design needs parallel copies with a join, a conditional early stop and a pause that can resume, and LangGraph provides each of those |
| LangChain's Ollama integration | langchain-ollama 1.1.0, with langchain-core 1.6.3 | ChatOllama talks to the local model, and with_structured_output returns a validated object | It is the one small part of LangChain the graph needs. Nothing else from LangChain is imported |
| Pydantic | 2.13.5 | Defines the Score and Scores shapes the model must return | The model's answer is checked against a schema before any code uses it |
Ollama and qwen2.5:14b | server 0.34.2 | Runs the model on this machine and scores the documents and writes the note | No API key, nothing leaves the machine, and the model was already installed. It returned valid structured output in every run I made |
| A2A Python SDK | a2a-sdk 1.1.4 | Reads the Agent Card, works out which transport to use and sends the message | It removes the hand-built JSON-RPC I used in part 2, which was there to show the wire format |
| httpx | 0.28.1 | Calls the federalregister.gov API for the verification step | Async HTTP, and it is already a dependency of the A2A SDK |
Not used: the prebuilt ReAct agent (create_react_agent, now create_agent in LangChain) and its tool-calling loop (the reasoning is under "Why the remote agent is called from a node" below), RemoteGraph and the LangGraph server, LangSmith, MCP, and the SDK's server side.
A2A does not require LangGraph. The protocol only defines what crosses the boundary between two agents, so the remote agent could be built on anything and my client could be a plain script. LangGraph is one way to write the client side.
Do you need LangGraph for this?
No. For three rows and one pause, the same steps fit in a plain async script: asyncio.gather for the verification, an if for the early stop and input() for the approval. That version would be shorter, and for a task this small it is arguably the better choice. LangGraph earns its place in three ways:
- The pause is a real stop. State is saved after every step, so interrupt ends the run and Command(resume=...) continues it, instead of a call blocked on the keyboard. In this script both halves run in one process, and MemorySaver keeps state in memory only, so I have shown the mechanism and not a run that survives a restart. That would need a persistent checkpointer.
- The graph can be drawn from the code. build_graph().get_graph() reported the nodes and edges, and I used them for the diagram below.
- Merging parallel results is declared, not written. The operator.add reducer says how the three verification results combine.
The cost is more concepts to learn and one more layer between you and the control flow. I judge the trade worth it here because this post is about orchestration, and LangGraph makes each part of it visible and named.
LangChain, LangGraph and Langflow
The three names look like one product and are not. LangChain and LangGraph are separate libraries from the same company, and one sits on top of the other. Langflow is a different kind of tool, mentioned here in passing because I use neither its canvas nor its runtime. None of the three is required to build an agent, as the last part of this section explains.
| LangChain | LangGraph | Langflow | |
|---|---|---|---|
| What it is | Building blocks for LLM applications: model integrations, tools, prompts and ready-made agents | A framework and runtime for stateful agents, written as a graph of nodes joined by edges | A visual tool for building AI agents and workflows on a canvas, with a no-code or low-code interface |
| You work in | Code | Code | A browser canvas |
| Who decides what runs next | Mostly the agent loop: the model picks a tool, reads the result and repeats | You, through the edges you draw, with agentic steps where you choose | The flow you connect on the canvas |
| Suits | Getting an agent running quickly | Precise control of every step, mixing fixed and model-driven steps, saved state and human approval | Prototyping multi-step or multi-agent applications quickly |
| In this post | Only ChatOllama and structured output | The whole control flow | Not used |
The LangChain documentation puts the relationship plainly: "LangChain's agents are built on top of LangGraph. This allows us to take advantage of LangGraph's durable execution, human-in-the-loop support, persistence, and more." It describes create_agent as "a minimal, highly configurable agent harness", and describes LangGraph as "a low-level orchestration framework and runtime for building, managing, and deploying long-running, stateful agents." It also says LangGraph "can be used without LangChain". So the choice is not one library against the other. It is how much of the control flow you want a prebuilt harness to decide for you.
The pre-built agent has moved between them. When I first ran a LangGraph ReAct agent for this project, the import of create_react_agent printed a deprecation warning saying it "has been moved to langchain.agents" and to use create_agent instead. The idea is unchanged: a model that chooses tools in a loop. I wrote about that loop in my 2024 post on LangChain Agents, where an LLM decides which tools to call, runs them and repeats until the task is done. LangGraph draws the same loop as a cycle between an agent node and a tools node, which is the shape I contrast with this graph below.
I chose LangGraph directly for this script because the orchestration is the point. A prebuilt harness makes the "what next" decision inside a model loop, and I wanted those decisions in edges I can read, test and draw. I still use one LangChain piece, ChatOllama, because it is a convenient way to call a local model and get structured output back. The LangChain documentation recommends its higher-level agents for people getting started, and LangGraph for "precise control over every part of your agent's behavior."
Langflow sits in a different place. Where LangGraph is a graph you write in Python and can print as Mermaid text, Langflow lets you connect models, tools, prompts and memory as boxes on a canvas. I tried it in October 2024 in my post Langflow: I installed it with Python 3.10, started it with python -m langflow run, opened its canvas at http://127.0.0.1:7860 and imported a "Doc to Podcast" flow, which needed an openai_api_key variable. That is a good way to prototype quickly. I have not used Langflow for this task, and I have not checked how its current version works, so treat this paragraph as a pointer, not a comparison I have tested. If you want to see what a flow looks like before writing code, start there. If you need the flow to be tested, versioned and driven from code, as an orchestrator that must verify a third party's output does, a code-level graph like LangGraph is the closer fit.
You do not need a framework to build an agent
None of the three is required. An LLM-powered agent is a loop, and you can write that loop yourself. Send the conversation and a list of tool definitions to the model. If the reply asks for tool calls, run them, add the results to the conversation and go round again. If it does not, the reply is the answer. Put a cap on the number of trips so a confused model cannot loop forever.
That is the whole idea, and I built it without an agent framework such as LangChain or LangGraph in Claude Code, part 9, where I used Claude Code to create a local coding agent. Its tech stack is Next.js, TypeScript, Tailwind and Ollama running qwen2.5-coder:7b, and its core is one function, runAgent. The function follows the diagram above. It calls the model with the conversation and the tool definitions, runs any requested tools and feeds the results back, up to 20 iterations (MAX_TOOL_ITERATIONS). The agent has four tools: read_file, write_file, run_command and search_files. The function reports progress through a callback, so the same loop serves both the browser interface and the command line.
Writing the loop yourself also means owning its awkward cases. That post handles a model that does not support native function calling and puts its tool call as JSON inside the text of its reply, by scanning the text for it. Part 2 shows the other side of A2A working without a framework too: the remote helloworld agent has no agent framework and no model, and its whole "agent" is one function that echoes the request.
A framework is a convenience, not a requirement. What LangGraph adds to a loop like this is what this script uses it for: state saved between steps, a pause that can resume, parallel branches with a defined merge, and a graph you can draw from the code. If you do not need those, the loop above is enough. The question is whether the extras are worth another layer to learn, which is the same question I asked of this script under "Do you need LangGraph for this?" above.
What an orchestrator is responsible for
A2A describes the remote agent's side in detail: the card, the skills, the task states. The client that coordinates one or more agents has duties of its own, and most of them are easy to lose when the model is left to improvise.
Step 5 deserves emphasis. The remote agent is a third party, and the official samples' own README says to treat everything an external agent returns, including its card and messages, as untrusted input. The verification step in this graph is that advice made concrete.
| Pattern | What it means | In this post |
|---|---|---|
| Sequential | Step B needs step A's output | Discover, fetch, score, draft |
| Parallel fan-out and fan-in | Independent calls run together and a join waits for all | One verification per document |
| Conditional routing | The next step depends on a result | Stop early if nothing is relevant |
| Supervisor | One agent routes work to several specialists | Not used |
| Handoff | Control moves to another agent for the rest of the conversation | Not used |
| Human in the loop | The run pauses for approval before an irreversible step | Approve the note before it can be sent |
The graph
The control flow is a LangGraph StateGraph. This diagram shows the nodes, the edge types, the shared state, the checkpointer and the pause.
| Part | What it is | Where it appears here |
|---|---|---|
| Node | A function that receives the State and returns updates to it | The six named boxes |
| Fixed edge | Always go from A to B | The thick arrows |
| Conditional edge | A function chooses the destination at run time, and can return several | fan_out and route |
Send | Start a node with its own private input, one copy per item | One verify_document per row |
| Reducer | A rule for merging updates to a field from parallel nodes | operator.add on checked |
| Checkpointer | Saves the State after each step so a run can stop and continue | MemorySaver, which makes the pause possible |
| Interrupt | A node stops the run and returns a value to the caller | human_review |
| Cycle | An edge back to an earlier node | None in this graph |
A ReAct agent, the pattern behind LangGraph's older create_react_agent (now moved to LangChain as create_agent), is the opposite shape: agent and tools are joined in a cycle and the model decides how many times around it goes. This graph has no cycle. Every node runs once, in an order I fixed. A cycle is one more edge if a loop is what you want, for example retrying a failed call or waiting while a remote task is still working, and nothing in this run needed one.
Why the remote agent is called from a node
fetch_briefing is an ordinary async function that calls the A2A client. The model never sees the remote agent, so it cannot skip the call, run it early or fill in its arguments. I chose this because fetching the briefing is always required and nothing about it needs judgment. The alternative is to expose the agent as a tool, which suits a supervisor that must pick among several agents, and the next subsection compares the two. I tried that pattern in an earlier draft of this post with a different agent and a small local model, and the model skipped a tool, called one before its inputs existed and passed wrong values. That is evidence from one model, and a stronger one might behave better, but a node removes the question.
The model has two jobs here, scoring relevance and writing a paragraph. It never sees the remote agent's text. It sees fields taken from the official federalregister.gov record, and code, not the model, adds the links to the note.
Remote agents and sub-agents as tools
Wrapping an agent as a tool is a documented and common pattern, and it is the main alternative to what this script does. LangChain's multi-agent guide lists it first, under the name subagents: "A main agent coordinates subagents as tools. All routing passes through the main agent, which decides when and how to invoke each subagent." It says the pattern suits parallel work and large contexts, because each sub-agent works "in isolation with only its relevant context." The guide does not mention remote agents or A2A, so it does not say whether a sub-agent runs in the same process or across a network.
In part 1 I measured this pattern against a router and against one big prompt for a different job, looking things up in six documents I had written. Subagents were no more accurate there than a router or a skills agent, and they used the most prompt tokens of the three designs that fit the model's window. That result is about documents you can load into one prompt. A remote agent is the case the pattern is really for, because you cannot load its data or its logic into your own context at all, and the boundary is real.
A sub-agent is an agent with its own loop and its own context, called by a main agent. From the main agent's side it is a name, a description and some arguments, and what comes back is a result. That shape is the same whether the sub-agent runs locally or is a remote A2A agent behind a wrapper. The network only adds what this post has been dealing with: finding the agent, deciding whether to trust its reply, and handling task states and rate limits.
| Graph node (this script) | Sub-agent as a tool | Handoff | |
|---|---|---|---|
| Who decides to call it | The graph, through its edges | The main agent's model | The current agent passes control, through a tool call |
| What the model sees | Nothing about the remote agent | A name, a description and arguments, then the result | The conversation moves to the other agent |
| Suits | A step that is always needed, checked before use, or run in parallel | A main agent that must choose among several agents, calling well-defined skills | A conversation the other agent should carry on |
| Main risk | The graph cannot adapt on its own to a case you did not draw | The model skips it, misorders it or fills in wrong arguments | The original agent loses control of the outcome |
The handoff column and the "main risk" row are my own summary, not quotations from the guide. The A2A documentation accepts tool-style wrapping and marks its limit. On its page about A2A and MCP, exposing an agent as a tool "works best when the skills are well-defined and can be called in a tool-like, stateless way", and "A2A's main strength is its support for flexible, stateful, collaborative interactions that go beyond a typical tool call." The two agents in this series are the tool-like kind: one well-defined request and one answer. Wrapping either as a tool would work.
Three points are my own inferences, not documented claims:
- A wrapper hides the parts of A2A beyond a tool call. A tool returns one result. Streaming, multi-turn exchanges and task states such as input-required have to be squeezed into that result, or into an error the model can read. fetch_briefing raises on failed, rejected, input-required and auth-required states, and a tool wrapper would have to decide what the model should be told instead.
- A wrapper decides what the model reads. If the remote agent's text goes straight back to the model, an unverified third party is writing into the model's context. This script filters the reply and verifies each row first, and a tool wrapper should do the same before returning anything.
- The patterns combine. The same guide says "You can mix patterns! For example, a subagents architecture can invoke tools that invoke custom workflows." A graph like this one could be a single tool of a larger supervisor.
My rule of thumb is to use a node when the call is always needed and nothing about it needs judgment, and a tool when the model genuinely has to choose which agent to call. My own experience is one small local model over a few runs, in which it skipped a tool and misordered calls. That is not a measurement across models, and a stronger model may need the safeguards less.
The code
regulatory_watch.py:
PYTHON
1import asyncio 2import json 3import operator 4import re 5import sys 6import uuid 7from typing import Annotated, TypedDict 8 9import httpx 10from a2a.client import A2ACardResolver, ClientConfig, create_client 11from a2a.types import Message, Part, Role, SendMessageRequest, TaskState 12from langchain_ollama import ChatOllama 13from langgraph.checkpoint.memory import MemorySaver 14from langgraph.graph import END, START, StateGraph 15from langgraph.types import Command, Send, interrupt 16from pydantic import BaseModel 17 18AGENT_URL = "https://agentnative.cazimedia.com" 19SKILL_ID = "federal-register-query" 20BRIEFING_PROMPT = "Give me a free current Federal Register briefing with provenance." 21OFFICIAL_API = "https://www.federalregister.gov/api/v1/documents/{number}.json" 22OFFICIAL_FIELDS = ["document_number", "title", "type", "publication_date", "html_url", "agencies", "abstract"] 23MODEL = "qwen2.5:14b" 24DEFAULT_ROLE = "export compliance officer at a manufacturer of underwater and defense electronics" 25RECIPIENT = "Dana Okafor, general counsel" 26 27BAD_STATES = {TaskState.TASK_STATE_FAILED, TaskState.TASK_STATE_REJECTED, 28 TaskState.TASK_STATE_INPUT_REQUIRED, TaskState.TASK_STATE_AUTH_REQUIRED} 29 30 31class State(TypedDict): 32 role: str 33 card: dict 34 rows: list[dict] 35 insights: dict 36 checked: Annotated[list[dict], operator.add] # parallel branches append here 37 picks: list[dict] 38 note: str 39 decision: str 40 41 42class Score(BaseModel): 43 document_number: str 44 relevance: int 45 reason: str 46 47 48class Scores(BaseModel): 49 scores: list[Score] 50 51 52def reply_text(event) -> str: 53 kind = event.WhichOneof("payload") 54 if kind == "message": 55 return "".join(p.text for p in event.message.parts) 56 if kind == "task": 57 if event.task.status.state in BAD_STATES: 58 raise RuntimeError(f"task ended {TaskState.Name(event.task.status.state)}") 59 return "".join(p.text for a in event.task.artifacts for p in a.parts) 60 return "" 61 62 63async def discover(state: State) -> dict: 64 async with httpx.AsyncClient() as http: 65 card = await A2ACardResolver(http, AGENT_URL).get_agent_card() 66 skills = [s.id for s in card.skills] 67 if SKILL_ID not in skills: 68 raise RuntimeError(f"{card.name} does not advertise {SKILL_ID}") 69 iface = card.supported_interfaces[0] 70 print(f" [discover] {card.name} v{card.version}: {iface.protocol_binding} " 71 f"{iface.protocol_version}, {len(skills)} skills, {SKILL_ID} found") 72 return {"card": {"name": card.name, "skills": skills}} 73 74 75async def fetch_briefing(state: State) -> dict: 76 client = await create_client(AGENT_URL, ClientConfig(streaming=False)) 77 text = "" 78 try: 79 message = Message(message_id=str(uuid.uuid4()), role=Role.ROLE_USER, 80 parts=[Part(text=BRIEFING_PROMPT)]) 81 async for event in client.send_message(SendMessageRequest(message=message)): 82 text += reply_text(event) 83 finally: 84 await client.close() 85 86 reply = json.loads(text) 87 if reply.get("type") != "federal-register-briefing" or reply.get("status") != 200: 88 raise RuntimeError(f"unexpected reply: {text[:200]}") 89 data = reply["data"] 90 91 # The agent is an unverified third party: keep only the fields this graph needs, 92 # in the shape it expects. The upgrade offers and everything else are dropped. 93 rows = [] 94 for r in data["rows"]: 95 if (re.fullmatch(r"\d{4}-\d{4,6}", str(r.get("document_number", ""))) 96 and str(r.get("url", "")).startswith("https://www.federalregister.gov/")): 97 rows.append({"document_number": r["document_number"], "title": str(r.get("title", ""))[:300], 98 "type": str(r.get("type", "")), "publication_date": str(r.get("publication_date", "")), 99 "url": r["url"]}) 100 ins = data.get("insights", {}) 101 insights = {"documents_analyzed": ins.get("documents_analyzed"), "document_mix": ins.get("document_mix"), 102 "top_agencies": ins.get("top_agencies", [])[:5]} 103 remaining = data.get("rate_limit", {}).get("remaining") 104 print(f" [fetch_briefing] {len(rows)} rows kept of {len(data['rows'])}; " 105 f"insights cover {insights['documents_analyzed']} documents; free-tier calls remaining: {remaining}") 106 return {"rows": rows, "insights": insights} 107 108 109def fan_out(state: State) -> list[Send]: 110 return [Send("verify_document", {"row": r}) for r in state["rows"]] 111 112 113def squash(s: str) -> str: 114 return re.sub(r"\s+", " ", s).strip() 115 116 117async def verify_document(payload: dict) -> dict: 118 row = payload["row"] 119 params = [("fields[]", f) for f in OFFICIAL_FIELDS] 120 async with httpx.AsyncClient(timeout=30) as http: 121 resp = await http.get(OFFICIAL_API.format(number=row["document_number"]), params=params) 122 problems, official = [], None 123 if resp.status_code == 404: 124 problems.append("not found at federalregister.gov") 125 else: 126 resp.raise_for_status() 127 official = resp.json() 128 for field, official_field in (("title", "title"), ("type", "type"), 129 ("publication_date", "publication_date"), ("url", "html_url")): 130 if squash(row[field]) != squash(official[official_field]): 131 problems.append(f"{field} differs from the official record") 132 ok = not problems 133 print(f" [verify_document] {row['document_number']}: {'matches the official record' if ok else '; '.join(problems)}") 134 result = {"document_number": row["document_number"], "ok": ok, "problems": problems} 135 if ok: 136 result["official"] = { 137 "title": official["title"], "type": official["type"], "date": official["publication_date"], 138 "agencies": [a["name"] for a in official["agencies"]], "url": official["html_url"], 139 "abstract": (official.get("abstract") or "")[:700]} 140 return {"checked": [result]} 141 142 143def score_relevance(state: State) -> dict: 144 good = {c["document_number"]: c["official"] for c in state["checked"] if c["ok"]} 145 if not good: 146 return {"picks": []} 147 docs = "\n\n".join( 148 f"{n}: {d['title']}\nType: {d['type']}. Agencies: {', '.join(d['agencies'])}. Published {d['date']}.\n" 149 f"Abstract: {d['abstract'] or '(none published)'}" for n, d in good.items()) 150 model = ChatOllama(model=MODEL, temperature=0).with_structured_output(Scores) 151 result = model.invoke( 152 f"I am a {state['role']}. Below are new Federal Register documents. For each one, give a " 153 "relevance score from 0 (irrelevant to my work) to 10 (I must read it today) and a one-sentence " 154 "reason that refers to my role. Use each document_number exactly as given.\n\n" + docs) 155 picks = [] 156 for s in result.scores: 157 if s.document_number in good and s.relevance >= 6: 158 picks.append({**good[s.document_number], "document_number": s.document_number, 159 "relevance": min(s.relevance, 10), "reason": s.reason}) 160 for s in result.scores: 161 print(f" [score_relevance] {s.document_number}: {s.relevance}/10 - {s.reason}") 162 return {"picks": sorted(picks, key=lambda p: -p["relevance"])} 163 164 165def route(state: State) -> str: 166 return "draft_note" if state["picks"] else "nothing_relevant" 167 168 169def draft_note(state: State) -> dict: 170 items = "\n".join(f"- {p['title']} ({p['type']}, {p['date']}). Why it matters: {p['reason']}" 171 for p in state["picks"]) 172 body = ChatOllama(model=MODEL, temperature=0).invoke( 173 f"Write the body of a short note (under 120 words) from Neil, a {state['role']}, to {RECIPIENT}, " 174 "flagging the Federal Register documents below. Do not include links, document numbers, a greeting " 175 f"or a sign-off, and do not add facts that are not listed.\n\n{items}").content.strip() 176 # Links come from the verified official records, never from the model. 177 sources = "\n".join(f"- {p['title']} ({p['type']}, {p['date']}): {p['url']}" for p in state["picks"]) 178 note = f"Dear {RECIPIENT.split(',')[0].split()[0]},\n\n{body}\n\nSources:\n{sources}\n\nBest regards,\nNeil" 179 print(" [draft_note] drafted") 180 return {"note": note} 181 182 183def human_review(state: State) -> dict: 184 return {"decision": interrupt({"note": state["note"], "question": "Send this note? (approve / reject)"})} 185 186 187def build_graph(): 188 g = StateGraph(State) 189 for name, fn in (("discover", discover), ("fetch_briefing", fetch_briefing), 190 ("verify_document", verify_document), ("score_relevance", score_relevance), 191 ("draft_note", draft_note), ("human_review", human_review)): 192 g.add_node(name, fn) 193 g.add_edge(START, "discover") 194 g.add_edge("discover", "fetch_briefing") 195 g.add_conditional_edges("fetch_briefing", fan_out, ["verify_document"]) 196 g.add_edge("verify_document", "score_relevance") 197 g.add_conditional_edges("score_relevance", route, {"draft_note": "draft_note", "nothing_relevant": END}) 198 g.add_edge("draft_note", "human_review") 199 g.add_edge("human_review", END) 200 return g.compile(checkpointer=MemorySaver()) 201 202 203async def run_until_pause(graph, payload, config): 204 async for update in graph.astream(payload, config, stream_mode="updates"): 205 if "__interrupt__" in update: 206 return update["__interrupt__"][0].value 207 return None 208 209 210async def tamper_test() -> None: 211 real = {"document_number": "2026-19211", "title": "International Traffic in Arms Regulations: " 212 "Modification of U.S. Munitions List Category XX(a)", "type": "Rule", 213 "publication_date": "2026-09-18", "url": "https://www.federalregister.gov/documents/2026/09/18/" 214 "2026-19211/international-traffic-in-arms-regulations-modification-of-us-munitions-list-category-xxa"} 215 for label, row in (("untouched row", real), 216 ("title altered", {**real, "title": "Repeal of all export controls"}), 217 ("invented document", {**real, "document_number": "2026-99999"})): 218 print(f"{label}:") 219 await verify_document({"row": row}) 220 221 222async def main(role: str) -> None: 223 graph = build_graph() 224 config = {"configurable": {"thread_id": "watch-1"}} 225 print(f"role: {role}\n--- run ---") 226 pending = await run_until_pause(graph, {"role": role, "checked": []}, config) 227 if pending is None: 228 print("\n--- finished without pausing: no document was relevant, nothing drafted ---") 229 return 230 print("\n--- paused for human review ---") 231 print(pending["note"]) 232 answer = input(f"\n{pending['question']} ").strip() or "reject" 233 print("\n--- resumed ---") 234 await run_until_pause(graph, Command(resume=answer), config) 235 final = (await graph.aget_state(config)).values 236 print(f"decision recorded: {final['decision']}") 237 238 239if __name__ == "__main__": 240 if "--tamper-test" in sys.argv: 241 asyncio.run(tamper_test()) 242 elif "--graph" in sys.argv: 243 print(build_graph().get_graph().draw_mermaid()) 244 else: 245 asyncio.run(main(next((a for a in sys.argv[1:] if not a.startswith("--")), DEFAULT_ROLE)))
Running it with uv
uv is a fast Python package and version manager. It can download a Python interpreter of its own, which matters here: the script needs Python 3.10 or newer, and the system Python on my Mac was 3.9.6, as I found in the first post. On macOS I installed it with Homebrew:
BASH
1brew install uv
Before you start
- Ollama must be running with the model pulled. The scoring step is the first place the script uses the model, so a missing model or a stopped server would show up there:
BASH
1 ollama pull qwen2.5:14b 2 curl http://localhost:11434/api/version
The curl prints Ollama's version if the server is up. The model is about 9 GB.
- Make a folder and save two files in it. Save the code above as regulatory_watch.py, and save this as requirements.txt:
CODE
1 langgraph 2 langchain-ollama 3 httpx 4 a2a-sdk
These are unpinned. The versions I ran are in the table above, and you can pin them, for example langgraph==1.2.11, if you want the same behavior later.
Option 1: a virtual environment
BASH
1uv venv --python 3.12 .venv 2source .venv/bin/activate 3uv pip install -r requirements.txt
uv venv --python 3.12 .venv creates an isolated environment in a folder called .venv, and downloads a CPython 3.12 build if the machine does not already have one. source .venv/bin/activate makes that environment the current one for this terminal window. uv pip install -r requirements.txt installs the four packages, and everything they depend on, into it. It uses uv's pip-compatible interface, so it reads an ordinary requirements.txt.
Then run the script. Start with the check that needs no model and no call to the remote agent, only the internet:
BASH
1python regulatory_watch.py --tamper-test
Then the graph itself:
BASH
1python regulatory_watch.py # the default role 2python regulatory_watch.py "marine biologist restoring coral reefs" 3echo approve | python regulatory_watch.py # answers the approval prompt for you 4python regulatory_watch.py --graph # prints the compiled graph as Mermaid text
The role is the first argument that does not start with --, so quote it if it has spaces. At the pause the script prints the draft note and waits: type approve or reject and press Return. The echo approve | form feeds that answer in, which is how I captured the run below. If nothing is relevant to the role, the run ends before the pause, as in my second run.
Option 2: uv run, with nothing to activate
BASH
1uv run --python 3.12 --with-requirements requirements.txt regulatory_watch.py --tamper-test
uv run builds a temporary environment from the requirements file, keeps it in uv's cache, runs the script inside it and leaves nothing installed on the rest of the machine. Everything after the script name is passed to the script, so the same arguments work:
BASH
1uv run --python 3.12 --with-requirements requirements.txt regulatory_watch.py "marine biologist restoring coral reefs"
I ran the --tamper-test command both ways in a clean folder and it produced the output shown below each time. uv also has a project workflow, uv init and uv add, which records dependencies in a pyproject.toml. I did not use it here.
What to expect
- The output will differ from mine. The briefing changes every day, and a model at temperature 0 is repeatable for a given input, not across different inputs.
- The remote agent is rate limited. Its reply reported a limit of 20 calls. If it stops returning a briefing, the script stops with an error in fetch_briefing instead of carrying on with bad data.
- The first run is slower. uv downloads Python and the packages, and Ollama loads the model into memory.
Runs
Before running the whole graph, I tested the verification step alone, since a check that has never failed proves little. I fed it one untouched row, one with an altered title and one invented document, using no model and no calls to the remote agent:
CODE
1untouched row: 2 [verify_document] 2026-19211: matches the official record 3title altered: 4 [verify_document] 2026-19211: title differs from the official record 5invented document: 6 [verify_document] 2026-99999: not found at federalregister.gov
Then the whole graph, with the role of an export compliance officer, typing approve at the pause:
CODE
1role: export compliance officer at a manufacturer of underwater and defense electronics 2--- run --- 3 [discover] AgentNative Data Exchange v0.3.0: JSONRPC 1.0, 6 skills, federal-register-query found 4 [fetch_briefing] 3 rows kept of 3; insights cover 25 documents; free-tier calls remaining: 19 5 [verify_document] 2026-19222: matches the official record 6 [verify_document] 2026-19211: matches the official record 7 [verify_document] 2026-19251: matches the official record 8 [score_relevance] 2026-19251: 0/10 - This document pertains to drug transit and production and is not directly related to underwater or defense electronics. 9 [score_relevance] 2026-19222: 0/10 - This document concerns employment regulations and does not pertain to export compliance for underwater or defense electronics. 10 [score_relevance] 2026-19211: 10/10 - This document directly affects the export regulations for uncrewed underwater vehicles, which is highly relevant to my role in export compliance for defense electronics. 11 [draft_note] drafted 12 13--- paused for human review --- 14Dear Dana, 15 16Please be advised of a recent publication in the Federal Register regarding the International Traffic in Arms Regulations: Modification of U.S. Munitions List Category XX(a). This modification is crucial as it directly impacts the export regulations for uncrewed underwater vehicles, an area highly relevant to our export compliance for defense electronics. I recommend we review the implications of this change on our current export practices and ensure compliance with the updated regulations. 17 18Sources: 19- International Traffic in Arms Regulations: Modification of U.S. Munitions List Category XX(a) (Rule, 2026-09-18): https://www.federalregister.gov/documents/2026/09/18/2026-19211/international-traffic-in-arms-regulations-modification-of-us-munitions-list-category-xxa 20 21Best regards, 22Neil 23 24Send this note? (approve / reject) approve 25--- resumed --- 26decision recorded: approve
The three verify_document lines run in parallel and finish in whatever order they finish: the order differed between my two full runs. All three rows matched the official record, so the model saw all three. It scored the two irrelevant documents 0 and the International Traffic in Arms Regulations rule 10, which fits the role, since the rule removes certain uncrewed underwater vehicles from the U.S. Munitions List. The note's links come from the verified record.
Then the same graph with a role that should match nothing:
CODE
1role: marine biologist restoring coral reefs 2--- run --- 3 [discover] AgentNative Data Exchange v0.3.0: JSONRPC 1.0, 6 skills, federal-register-query found 4 [fetch_briefing] 3 rows kept of 3; insights cover 25 documents; free-tier calls remaining: 19 5 [verify_document] 2026-19211: matches the official record 6 [verify_document] 2026-19251: matches the official record 7 [verify_document] 2026-19222: matches the official record 8 [score_relevance] 2026-19251: 0/10 - This document is about drug transit and production and has no direct relation to marine biology or coral reef restoration. 9 [score_relevance] 2026-19222: 0/10 - This document pertains to employment regulations and does not impact marine biology or coral reef restoration efforts. 10 [score_relevance] 2026-19211: 2/10 - This document modifies regulations on uncrewed underwater vehicles, which could indirectly affect marine research and conservation efforts, but is not directly relevant to coral reef restoration. 11 12--- finished without pausing: no document was relevant, nothing drafted ---
This time the conditional edge sent the run to __end__ before drafting, so nothing was written and there was nothing to approve. The model gave the ITAR rule 2 out of 10 for a marine biologist, below the threshold of 6 that I set in code.
What this does and does not show
- The agent adds little for this data. The Federal Register publishes a free, keyless API, which is what the verification step calls. The three free rows are ones I could have fetched from it directly. The agent's value is in its paid tier of filters, up to 100 rows and aggregations, which I did not use. What the demo does show is the orchestration pattern: discover, delegate, verify, merge, hand back to a person.
- Reply handling was only partly exercised. The agent answered with a plain message, so the task-state branch (failed, rejected, input-required, auth-required) never ran. The agent also advertises a durable task for making datasets query-ready, which I did not try.
- No contextId was sent. Nothing here needs conversation state. The LangChain A2A documentation describes how an Agent Server maps a contextId to a LangGraph thread_id, which matters for agents that remember earlier turns. The thread_id in my checkpointer is a separate, local identifier.
- RemoteGraph was not used. It calls a graph deployed on a LangGraph server over that server's own API, and this agent is not one.
- The scoring rests on three documents. A model choosing among three rows is a thin test, and its scores are one small local model's opinion. I set the threshold of 6 by hand.
- The listing is unclaimed. Its owner has not verified it with the registry, so I treated its replies as untrusted and ignored the upgrade offers.
- The free tier is limited. Its own rate_limit field reported a limit of 20 calls, and I made a handful of calls in total.
- My earlier posts: LangChain Agents, Langflow and Claude Code, part 9, an agent loop with no framework
- LangChain: A2A endpoint in Agent Server, including contextId and thread_id