Logistics / B2B SaaSInteractive Prototype

Turning manual freight invoice review into automated audit intelligence.

€1M+ in freight savings · 60% faster invoice validation · 45% increase in weekly engagement

Freight teams processing hundreds of invoices a month have no reliable way to know when they’re being overbilled. Manual review is slow, inconsistent, and misses discrepancies that compound over time — by the time an overcharge pattern is spotted, months of overpayment have already cleared.

This is an anonymized reconstruction of real product work. The workflow UI, data, names, and examples are synthetic; the stated outcomes reflect the original programme.

The problem

Every overcharge that is not caught is paid.

Freight carriers invoice against complex rate structures — base freight, surcharges, handling fees, fuel adjustments — each governed by contract terms that change by lane, season, and carrier.

Manual review means:

  • Reviewers checking invoices against PDF contracts line by line
  • Discrepancies missed because the volume is too high to review thoroughly
  • No systematic record of overcharges by carrier or lane
  • Finance approving invoices on trust, not verification

The product opportunity was to replace that manual process with a workflow where AI surfaces the discrepancies and humans make the decisions — at a fraction of the time and with a full audit trail.

Prototype

Freight Invoice Audit Workflow

Load the sample invoice, review the AI findings on flagged line items, and record your decisions. All client-side with synthetic data.

Freight Invoice Audit

Drop a freight invoice to audit

PDF, CSV or structured data — or use the sample below

Key product decisions

Four decisions that shaped the workflow.

  1. 01

    AI flags, humans decide — always.

    Freight invoices involve contractual and commercial relationships. Automatically rejecting a line item without human review would damage carrier relationships and create dispute risk. AI surfaces the discrepancy; the reviewer owns the decision.

  2. 02

    Show the contract reference, not just the flag.

    A flag without evidence is noise. Every AI finding links to the specific contract clause it references — so reviewers can validate the finding in seconds rather than hunting through contract PDFs.

  3. 03

    Line-item granularity over invoice-level summary.

    An invoice-level 'overpayment detected' gives you a number but not an action. Reviewers need to know which line item, what the discrepancy is, and what the contract says — at that level, every item is actionable.

  4. 04

    Audit trail by default, not on request.

    Freight audit creates dispute records. Every review decision — accept, dispute, override — needs to be attributable and timestamped. Building the audit trail as infrastructure from day one prevented a major rework when compliance asked for it.

What I chose not to build

Constraints that kept the product trustworthy.

Speed is not the goal when the product touches financial disputes and carrier relationships.

  1. 01

    No automated dispute submission

    Automatically submitting disputes to carriers would have created legal and commercial risk before humans reviewed the evidence. The workflow stops at a decision record — submission is always manual.

  2. 02

    No carrier portal integration in v1

    We scoped carrier API integration for v2. Building it in v1 would have delayed the core audit workflow by months and made the carrier-specific integration the critical path — not the product value.

  3. 03

    No ML model training on proprietary contract data

    Contract terms are sensitive and customer-specific. Training a shared model on customer contract data raised data-privacy concerns that weren't worth the marginal accuracy gain over deterministic rule-matching.

Risks & controls

Where the system has to be right.

Potential risks

  • False positive flags damaging carrier relationships
  • Incorrect contract reference causing wrongful disputes
  • Reviewer fatigue leading to bulk acceptance without review
  • Data model not handling multi-currency invoice structures
  • AI confidence on non-standard surcharge types

Controls

  • Every flag links to explicit contract clause — reviewers can verify
  • Confidence indicator shown per finding — low-confidence flagged separately
  • Bulk acceptance disabled — each flagged item requires individual decision
  • Multi-currency normalization before comparison logic runs
  • Unknown surcharge types escalated for manual review, never auto-flagged
  • Full audit log of every reviewer decision with timestamp and user
How we measured success

The measurement definitions we used.

Each metric below is defined by its measurement method — what was counted, how, and against what baseline.

  • Invoice validation time — median time from invoice intake to reviewer decision, compared against the pre-launch manual baseline
  • Flag accuracy rate — proportion of AI-flagged items where the dispute was upheld by the carrier
  • Reviewer decision time — median time per flagged line item from view to accept or dispute action
  • Dispute recovery rate — total amount recovered via upheld disputes as a proportion of total amount disputed
  • Reviewer override rate — proportion of flagged items accepted without dispute, indicating possible false positives
  • False positive rate — flagged items later accepted, segmented by surcharge type
  • Weekly active reviewers — count of unique reviewers completing at least one full invoice audit
Rollout

Prove accuracy before adding intelligence.

  1. Phase 1

    Rule-based audit

    • — Contract upload and parsing
    • — Deterministic discrepancy detection
    • — Reviewer queue
    • — Decision audit trail
  2. Phase 2

    AI-assisted review

    • — LLM-assisted flag reasoning
    • — Contract clause cross-reference
    • — Confidence scoring
  3. Phase 3

    Network intelligence

    • — Carrier dispute pattern analysis
    • — Benchmark overcharge rates by lane
    • — Proactive contract risk alerts
My contribution

Where I was involved.

  • Led product discovery including workflow mapping with freight audit operations teams and finance stakeholders
  • Defined the discrepancy detection logic and the boundary between AI-flagged and deterministic rule-based findings
  • Designed the reviewer queue, decision model, and audit trail specification
  • Set product requirements for contract parsing, multi-currency handling, and evidence linkage
  • Partnered with engineering and operations to deliver 60% faster invoice validation and €1M+ in recovered freight costs
  • Prioritised iteration based on reviewer feedback and false-positive rate data

Cross-functional collaboration

  • Product
  • Engineering
  • Freight Operations
  • Finance
  • Design
Retrospective

What I would do differently.

  • 01I would have run a structured false-positive calibration session with operations reviewers in week two — we learned too late that certain surcharge types had legitimate exceptions that the rule logic wasn't accounting for.
  • 02The contract parsing module was scoped as engineering work without enough PM involvement. In hindsight, I should have owned the contract data model more closely — several edge cases in complex contracts caused rework that better upfront discovery would have caught.
  • 03I'd define the dispute recovery metric from day one and instrument it in the product, not in a separate spreadsheet. Having it outside the product made it harder to tie product changes to outcome improvements.