The graph.
Preparer and Reviewer are parallel entry nodes of a GraphBuilder graph; Strands hands each only the rows. Both call the real line_guidance tool and answer with pydantic structured output. The Referee hangs off a conditional edge and runs only on a disagreement.
Bank export
Date, description, amount. Batches of twelve go into the graph.
Preparer
Posts each row to a Part I line and quotes the IRS instruction through the line_guidance tool.
Reviewer
Same job, never sees the Preparer. Agreement posts the row. Disagreement escalates.
Referee
Runs only on a split. Reads both quoted rules, decides, and its reason is kept with the row.
Foot and fill
Sums lines 9, 17, 18, word-checks every quoted rule, the Referee's included, against the IRS text, fills the PDF, stamps DRAFT.
Architecture diagram and source: docs/architecture.png, the repository. Model host: Google Gemini through the Strands GeminiModel, rotating model ids when a free-tier cap is hit. Amazon Bedrock is a one-line swap.
The recorded run, batch by batch.
| Batch | Rows | Model | Preparer | Reviewer | Referee | Tool calls | Wall |
|---|
Footing checked against filed returns.
The same formmath.py that fills the form was run over an IRS e-file XML batch, rebuilding each return's stated totals from its own line items. Reproduce with PYTHONPATH=src python scripts/validate.py.
| Line | Checked | Matched | Rate |
|---|