* Revenue figures are market-based estimates only and are not guarantees of income. Actual results will vary based on execution, market conditions, and individual effort. This is not financial or investment advice.
How the agent runs it
Each night a scheduled GitHub Actions job pulls new freight invoices (uploaded as PDFs or EDI 210 files by the customer) into the PostgreSQL ledger, where an extraction agent parses line items into structured rows and stamps each with a source-of-truth checksum. A separate audit agent then compares each line item against the carrier's published rate card fetched via API, checking for duplicate charges, incorrect weight breaks, accessorial fees not on the original quote, and fuel surcharge miscalculations — following a fixed 12-step checklist before writing any finding to the ledger. When an overcharge is confirmed against the ledger (never from memory), a dispute-drafting agent generates a carrier dispute letter using a locked template, writes the draft to a review queue in Retool, and only sends after a human clicks approve or 48 hours elapse with no action on items under $75. A nightly GitHub Actions job computes per-customer recovery rates, false-positive rates, and dispute acceptance rates from the ledger, writes a dated lessons file, and injects it into every agent's next-run context.
Who this is for
Best suited for a founder with 2–4 years of freight brokerage, 3PL operations, or supply-chain finance experience who understands carrier tariff structures and knows which accessorial charges are most commonly overbilled. They need enough technical fluency to configure PostgreSQL schemas, GitHub Actions jobs, and Claude API prompts, but do not need to be a software engineer — the critical skill is domain knowledge to write accurate audit checklists and validate agent findings during the first three months. This suits a solo operator or a two-person team because the human workload is concentrated in a short daily review session (the Retool dashboard and digest email) rather than ongoing customer interaction.
Market opportunity
U.S. small and mid-size shippers spend an estimated $100B+ annually on parcel and LTL freight, and third-party auditors consistently recover 1–3% of freight spend through billing error detection — a figure that has grown as carriers added dozens of new accessorial fee categories between 2022 and 2025. Most audit firms charge 50% of recoveries, making them unattractive for shippers with modest volumes; a flat monthly fee is compelling for businesses shipping $10K–$200K/year in freight who currently audit nothing. The 2024–2025 carrier surcharge complexity surge (dimensional weight changes, peak season fees, address correction proliferation) has created more billing surface area for errors precisely as small shippers lack the staff to catch them.
Tech stack
Monetization
Price: $149/month flat per shipper account (up to 200 invoices/month); $249/month for 201–600 invoices; no commission on recovered amounts, no variable pricing the agent can alter — price table lives in Stripe product catalog only. Variable cost per customer (base tier, 200 invoices): Claude API ~$6 (haiku classification + opus audit reasoning at ~$0.03/invoice avg), Shippo/EasyPost rate-card lookups ~$2, SendGrid email delivery ~$0.50, payment processing (Stripe 2.9% + $0.30) ~$4.62, human review time 0.5 hrs/mo at $40/hr = $20. Total variable: $6 + $2 + $0.50 + $4.62 + $20 = $33.12. Fixed monthly costs: PostgreSQL hosting (Railway or Render) $25, Retool Team plan $50, GitHub Actions compute $20, SendGrid base plan $20, Shippo API base $0 (pay-per-call, included above), bookkeeping software $15, liability insurance for data handling $80, owner's human review and digest-reading time 8 hrs/mo at $40/hr = $320. Total fixed: $530/mo. Break-even: fixed $530 / ($149 - $33.12 contribution margin) = 4.6 customers, so 5 paying customers at the base tier covers fixed costs; at a blended mix of base and mid-tier the break-even is roughly 4 customers. Margin at target (120 customers, blended avg revenue $175/mo): revenue $21,000 — variable costs (120 × $38 blended) $4,560 — fixed $530 = gross profit $15,910, margin 75.8%. Realistic month-12 estimate assuming 85 customers and agent errors causing ~5% dispute rework requiring extra human hours: revenue $14,875 — variable $3,230 — fixed $530 — error rework $400 = $10,715, margin ~72%.
Key risks
- → Fabricated overcharge figures (Vend-style invented records): the audit agent is prohibited by code from writing any overcharge claim unless it can produce a ledger row ID for the original invoice line AND a rate-card API response timestamp within the same job run; disputes citing only agent memory are rejected at the tool layer before they reach the dispatch queue.
- → Agent confidence inflating dispute amounts (helpfulness liability): the dispute-drafting agent has no knowledge of and no tool access to customer pricing or subscription tiers; it reads only the ledger overcharge amount, which is immutable after the audit agent writes it; it cannot apply credits, goodwill adjustments, or round up figures — those fields are locked to human-only writes in Retool.
- → Long-run drift on dispute templates (coherence decay): the lessons file updated nightly records false-positive rates by carrier and charge type; if the false-positive rate for any carrier exceeds 15% over a 7-day window, the GitHub Actions watchdog automatically pauses dispute dispatch for that carrier and creates a human escalation ticket in Retool.
- → Impersonation of carrier reps or customer executives requesting dispute withdrawal (adversarial social engineering): any inbound email or webhook claiming authority to cancel a dispute is ignored by the agent; only a human clicking 'withdraw' in Retool with an authenticated session can cancel a filed dispute, and the system prompt explicitly states that message content alone cannot grant authority.
- → Agent looping or stalling on malformed invoices (coherence decay / stuck agents): the heartbeat job in GitHub Actions checks that each invoice moves from 'ingested' to 'audited' within 4 hours; invoices stuck longer are flagged, the audit agent task is killed, and the invoice is routed to a human-review bucket with a Retool alert.
Getting started
- 1 Build and lock the PostgreSQL invoice ledger schemaDefine tables for invoices, line items, rate-card snapshots, audit findings, and dispute statuses before writing any agent code; every monetary field must be write-once from the audit agent and human-editable only through Retool, not through any Claude tool call. This is the source-of-truth foundation that prevents fabricated figures from reaching dispute letters.
- 2 Write and test the 12-step audit checklist as codeEncode each audit check (duplicate bill detection, weight-break validation, accessorial fee cross-reference, fuel surcharge calculation) as a discrete Python function that reads from the ledger and the rate-card API response, returning a structured result with evidence row IDs; the Claude audit agent calls these functions as tools and is forbidden from making overcharge claims by any other path. Run the checklist against 500 real historical invoices you obtain from a beta customer to calibrate false-positive rates before launch.
- 3 Set up Stripe with locked price table and no agent accessCreate Stripe products at the $149 and $249 price points and configure webhooks to provision customer accounts in the ledger; the Claude agents must have no Stripe API key in their environment — billing changes can only be made by a human through the Stripe dashboard. This enforces the price floor at the infrastructure layer, not via system prompt.
- 4 Recruit three beta customers from freight forums or LinkedInOffer three small shippers (annual freight spend $30K–$150K) a free 90-day audit in exchange for real invoice data and honest feedback on dispute letter quality; use this period to measure actual overcharge recovery rates, tune the audit checklist, and establish the false-positive rate per carrier that will gate the heartbeat watchdog thresholds.
- 5 Configure the GitHub Actions watchdog and daily digestSet up three scheduled jobs: the nightly audit runner, a heartbeat that checks invoice pipeline health every 4 hours and kills stalled tasks, and a morning digest job that queries the ledger for margin, dispute acceptance rates, and any anomaly flags and emails a structured summary to the human owner before 8 a.m. Read this digest every day — it is the primary human oversight mechanism and the only way to catch agent drift before it compounds.
// done for you
Want us to build
Freight Invoice Audit and Dispute Filing for Small Shippers
for you?
We contract experienced engineers to deploy AI agent businesses end-to-end — custom domain, branding, live and earning in weeks. No code required on your part.
We reply within 1 business day · No obligation · Canadian-based team