Automic BTRA

Automic · BTRA MVP

Banking Transaction Reconciliation Automation

Give the system a bank account extract and the matching Invana cash-transactions extract (the two tabs of the T+1 Bank Rec workbook). It runs a five-stage pipeline and, for every bank line, tries three matching tiers in order - cheapest and most reliable first - stopping at the first confident match. Whatever the rules and ML can't confidently decide is escalated; anything still unresolved is left unmatched for a person. It never posts to a ledger.

1 · Extract

Read both tabs (.xlsm or CSV)

→

2 · Label

Row roles: keep TXN rows

→

3 · Normalize

Canonical shape + sign

→

4 · Match

Manual → ML → AI

→

5 · Validate

vs the macro's output

Confidence bands: ≥0.95 auto-match · 0.75–0.95 review (manual queue) · <0.75 unmatched. Every match records the tier, rule, confidence and reason for full explainability.

1 · Extract

Reads the Bank Account Extract and Invana Data Extract straight from the .xlsm tabs (via openpyxl), or from the exported CSVs — keeping the raw report layout.

2 · Label (row roles)

Each row is tagged with a role so only real transactions are matched — report titles, headers, blank bands and totals are set aside.

META HEADER BLANK TOTAL TXN ✓

3 · Normalize

Both sides become one canonical shape — date, signed amount, direction, description — using the macro's sign convention so a bank debit lines up with the right Invana movement.

bank debit  → −abs(debit)
bank credit → +abs(credit)
   ↕ matched on ROUND(amount, 2) vs
Invana signed Movement Amount

Stage 4 · Match — three tiers, first confident hit wins

Deterministic first, intelligent second (design doc §05 / §08).

Tier 1 - Manual (rules)

the BuildBankRec macro

Reproduces the workbook's own VBA logic exactly — this is today's process, made explicit and testable.

Why first: zero-cost, deterministic and audited — it matches ~75% of lines the way the finance team already does.

Tier 2 - ML

btra-sklearn-logreg-match-v0.1

A trained match classifier that scores a bank line against each remaining candidate and returns a real match probability.

Backends: prefers scikit-learn; falls back to a self-contained numpy model, then fixed weights. Runs on what the rules leave over.

Tier 3 - AI

amazon-bedrock · claude-sonnet-4-5

Amazon Bedrock via the Converse API disambiguates the genuinely ambiguous remainder — several candidates competing for the same amount.

Graceful degradation: if credentials/model access are missing, a circuit breaker stops after a few errors and the run completes on rules + ML.

How the ML tier decides: features → probability

A model can't compare two transactions directly; it needs numbers. For each bank line ↔ candidate pair we compute six similarity features, each in 0–1. A logistic-regression model — trained to tell real matches from non-matches — turns those six numbers into a single match probability, which becomes the tier's confidence.

The six features

  • amount — how close the absolute amounts are
  • date — how close the dates are (decays over a week)
  • reference — shared reference/ID tokens
  • merchant / description — shared narrative words
  • historical — has this pairing matched before

Why logistic regression

It's the simplest model that outputs a calibrated probability (not just yes/no), trains in milliseconds with no GPU, and stays fully explainable — each feature has a weight you can read. It's the design doc's “replace hand-tuned weights with a trained ranking model” step, kept MVP-simple.

The model, in one line

p(match) = sigmoid( w · [amount, date, reference, merchant, description, historical] + b )

When the amounts and dates agree strongly, p is high; when they disagree, p collapses. Real bank narratives rarely share text with Invana descriptions, so amount + date proximity carry most of the signal — exactly why the tricky, same-amount-many-candidates cases are handed on to the AI tier.

1 · Featurize

Turn a pair into six 0–1 similarities.

2 · Score

Logistic regression → match probability.

3 · Band

≥0.95 auto · 0.75–0.95 review · <0.75 drop.

Stage 5 · Validate against the macro (ground truth)

The workbook's Bank Rec sheet is the macro's own output. The MVP compares its matched Cash Security per bank row against that sheet, so we can prove the rules tier reproduces today's process and measure what ML adds on top.

85.5%

agreement with the macro (4,421 / 5,170 rows)

0

cash-security disagreements where both matched

386

rows recovered beyond the macro by ML

363

macro “matches” that reused one Invana row → now review items

Reading it: where both matched, the MVP picked the same Cash Security every time (zero disagreements) — the rules tier is faithful. The only divergence is deliberate: the macro can reuse one Invana movement for many bank lines, while the MVP assigns each Invana row once, so a few “matches” become review items instead. Meanwhile ML recovers hundreds of rows the macro left as “No match”.

Explainability by design

Every matched group records its tier, rule, confidence, status, the reason, the matched Cash Security / Issuer and the ML features. Everything the rules and ML can't confidently place is a review item; the AI advises, a person confirms — nothing is written to a ledger.

auto
≥0.95, high-confidence match
review ⚑
0.75–0.95, manual queue
ml / ai
which tier decided it
unmatched
left for more evidence

Outputs & how to run

Each run writes a full explainability JSON and a self-contained HTML report (pipeline, summary + validation cards, matched groups, unmatched on both sides).

Full pipeline (rules + ML + AI), validated

./reconcile.sh

Offline (rules + ML, no Bedrock)

./reconcile.sh --no-ai

Macro rules only (reproduces the workbook)

./reconcile.sh --no-ml --no-ai

End-to-end flow

Bank + Invana extracts → Extract → Label → Normalize → Manual → ML → AI (Bedrock) → JSON + HTML + validation

For each bank line the first tier at or above its confidence band wins; otherwise the next tier is tried; otherwise the line stays unmatched for review.