Round Raises $6M to Build the AI-Powered Finance Automation Platform for Modern Finance Teams

How to evaluate an AI cash reconciliation agent before month-end in 2026

Author
Pac O'Shea
Date
20 September 2026
Reading time
7 min
Share

A practical test for AI cash reconciliation: known-answer data, false matches, exception handling, evidence, access and month-end sign-off.

Evaluate an AI cash reconciliation agent on your own known-answer data, not its best demo. It should prove why each match was made, isolate uncertainty, preserve every exception and produce a month-end record another person can reconstruct. The buying decision turns on false matches, not headline automation. A missed match creates visible work. A wrong match can make incomplete accounts look finished.

TL;DR

  • Build a known-answer test set from your own accounts and close patterns.
  • Score false matches separately from missed matches and unresolved items.
  • Require field-level reasons, source lineage and competing candidates for every match.
  • Send ambiguity to an owned exception queue with ageing and close evidence.
  • Accept the pilot only if its output supports period-end review without reconstruction.

What exactly are you evaluating?

Cash reconciliation is the controlled process of linking bank movements to accounting records, explaining differences and confirming what remains unresolved for a stated period. It is not the same job as consolidating balances into a cash view. Visibility tells you what appeared across accounts. Reconciliation proves how individual movements relate to the books.

That boundary matters because the word agent can hide several different products. One may propose matches. Another may create ledger entries. Another may simply summarise a queue. Write the operating job before comparing them.

Scope boundary: this evaluation covers transaction matching, exception handling and month-end evidence. It does not assess authority to move money, approve expenditure or release a transaction.

Write the job in one sentence: for each imported cash movement, find the supported accounting record or create an owned exception, then preserve enough evidence for period-end review. Anything beyond that sentence needs its own control decision. Anything short of it is assistance, not completed reconciliation.

Define the job before measuring the product
StageRequired outputControl question
IngestComplete, dated source recordsCan every line be traced to its origin?
ProposeCandidate match with reasonsWhich fields and rule produced it?
DecideAccepted match or visible exceptionWhat happens when evidence conflicts?
CloseReviewed period with open itemsCan another person reproduce the result?

Existing accounting systems already show why this is more than a similarity score. Oracle's NetSuite rules combine transaction number, amount and date, and stop automatic matching when two candidates remain equally plausible. Microsoft's advanced bank reconciliation also separates statement import, rule definition, automatic matching and later review. An AI product still has to meet those operational needs even if its technique is different.

Each proposed cash match carries source identifiers, decisive fields, alternatives and a route to either acceptance or review.

Build a known-answer test set before the demo

A useful pilot starts with records whose correct treatment is already known. Pull a representative slice from closed periods, preserve the source exports and label the expected result before the vendor runs anything. Do not let the product's output become the answer key.

Your set should include ordinary matches and the awkward cases that consume close time:

  • Exact matches with stable identifiers, amounts and dates.
  • Date shifts caused by weekends, settlement delay or posting cut-offs.
  • Grouped matches across one-to-many, many-to-one and many-to-many records.
  • Duplicate-looking lines where two candidates share an amount or reference.
  • Partial records with missing references or truncated descriptions.
  • True exceptions that should remain unmatched until a person resolves them.

Keep some cases hidden until the final run. A vendor that has tuned against every labelled example has demonstrated configuration effort, not general performance. The UK government's AI procurement guidance recommends defining acceptable performance, testing across conditions and preserving end-to-end auditability. Those principles fit reconciliation well.

Score the errors that matter to finance

One accuracy figure hides the main risk. Separate correct automatic matches, wrong automatic matches, correct suggestions sent for review, genuine matches the system missed and true exceptions it left open.

Example known-answer scorecard for 1,000 source lines
OutcomeCountWhat it tells you
Correct automatic matches880Work removed without changing the answer
Wrong automatic matches20Hidden correction risk created by the product
Correct suggestions for review15Useful assistance without silent acceptance
Genuine matches missed5Extra work left for the team
True exceptions left open80Correctly bounded unresolved work

In this test, 900 lines genuinely had a match. The product automatically accepted 900, but 20 of those were wrong. It found another 15 as reviewable suggestions and missed five. The operational question is not whether it touched most lines. It is whether the team would accept 20 hidden errors to remove 880 routine decisions.

Report results by entity, account, bank format, transaction type and source freshness. A single aggregate can look strong while one important account performs badly. Set separate acceptance limits for false matches, missed matches and unresolved exceptions, then agree what causes the pilot to stop.

A reconciliation scorecard keeps correct automatic matches, wrong automatic matches, review suggestions, missed matches and true exceptions separate.

Demand an explanation for every accepted match

Confidence should be an output, not an answer. Ask what evidence moved the result above the acceptance threshold and whether any competing candidate was close. A label such as high confidence is not useful without its basis.

For every accepted match, require a compact evidence record:

  1. Source identity: bank line and accounting record identifiers.
  2. Decisive fields: the references, amounts, dates and parties that agreed.
  3. Transformation: any normalisation, grouping or date tolerance applied.
  4. Alternative candidates: what else was considered and why it lost.
  5. Decision route: rule, model or human action that accepted the match.
  6. Version and time: the configuration used and when the decision happened.

This is not a demand for a technical account of every model weight. It is a demand for an operational reason a finance reviewer can test. NIST describes trustworthy AI as including valid and reliable, accountable and transparent, explainable and interpretable characteristics, applied through use, test and evaluation. Its AI Risk Management Framework FAQ also cautions that the characteristics work together and do not create trust on their own.

Make the exception queue part of the product test

A reconciliation agent should know when not to decide. If two candidates are equally plausible, the correct output is an exception with enough context for a person to choose. Oracle documents this exact boundary in its system rules: unresolved equal candidates require manual selection.

Test the queue as seriously as matching. Every item needs a stable ID, reason code, affected period, owner, age, next action and final close state. Grouped matches should show every component and the basis on which the group balances.

Exception states that support a controlled close
StateMeaningEvidence to retain
UnmatchedNo supported candidate existsSearch scope and fields tested
AmbiguousMore than one candidate remainsCandidate set and distinguishing fields
Data issueSource is missing, stale or malformedSource status and remediation owner
ResolvedA person or later record settled the itemActor, reason, time and linked record
Carried forwardThe item remains open by policyNamed owner and next review date

NetSuite's bank data guidance keeps matching, manual exceptions, submission and statement reconciliation as visible stages. That is a useful buyer test: does the new product preserve those distinctions, or collapse everything into a green completion count?

Test the month-end evidence, not only the live queue

The pilot is incomplete until a reviewer can close a period. Ask the vendor to produce the evidence pack from the same test run, without a spreadsheet assembled afterwards.

  • A source completeness check for every account and statement period.
  • A list of accepted matches with their evidence and decision route.
  • An open-item report grouped by reason, age, owner and materiality policy.
  • A history of manual changes, reopened items and configuration changes.
  • A reviewer sign-off tied to the exact version of the results.
  • An export that preserves stable identifiers for later audit sampling.

Your agent does not need to copy that layout, but it should leave finance with the same basic ability to see what tied out, what did not and which period was reviewed.

Then change something. Correct a source line, reopen a match and adjust a rule. The evidence should show what changed, who did it and which later output replaced the earlier one. A static export that cannot represent correction is not a durable close record.

A period closes only when source completeness, accepted matches, owned exceptions, change history and reviewer sign-off meet in one evidence record.

Use a staged pilot with explicit stop conditions

Run the product in observation mode first. Compare its output with the established close, review every proposed auto-match and record the time spent resolving exceptions. Only widen scope after the team understands where it fails.

  1. Define. Fix the accounts, period, data sources, expected outputs and acceptance limits.
  2. Back-test. Run the labelled set and score every error category separately.
  3. Shadow. Process a live period without letting the product's decisions replace the existing control.
  4. Bound. Allow automatic acceptance only for proven transaction classes and source conditions.
  5. Monitor. Sample accepted matches, review exception ageing and retest after configuration changes.

Stop or narrow the pilot when a wrong automatic match crosses the agreed limit, source lineage is lost, ambiguous items are silently selected, an evidence export cannot be reproduced or the queue has no accountable owner. These are product-control failures, not minor demo defects.

Where does Round fit today?

Our current Xero integration page says Round account feeds and invoices flow into Xero and are ready to reconcile. Our NetSuite integration page says records stay current and each Round transaction is reflected in NetSuite. The current pricing page lists Xero integration and places NetSuite plus custom ERP integrations in the Enterprise plan.

Those are accounting-data and integration claims. Our public pages checked on 20 September 2026 do not name a standalone AI cash reconciliation agent, publish match-confidence thresholds or specify the exception and month-end evidence fields in this guide. We would therefore ask the same questions of our own setup: which records move, which system performs the match, what remains manual, which plan applies and what evidence the connected system retains.

For the broader automation context, read how AI-orchestrated treasury works. For the separate job of consolidating balances, see real-time cash visibility across banks. Neither replaces the reconciliation evaluation described here.

Sources

Nothing within this blog is intended to be a recommendation. Round does not offer financial advice.

Frequently Asked Questions

It should prove that it finds genuine matches, avoids false matches, isolates uncertain items, explains the evidence behind each decision and produces a complete period-end record. Test those outcomes on a known-answer dataset from your own accounts.

A false automatic match is usually the most dangerous because it can make an account look complete while hiding the wrong pairing. Missed matches create work, but visible work is easier to control than a confidently wrong result.

Not because of the label alone. Ask what evidence produced the confidence, whether competing candidates existed and whether the same threshold has held on your data. Auto-acceptance should be bounded by transaction type, source quality and an agreed exception policy.

Keep the source line, matched record identifiers, rule or model route, decisive fields, confidence output, alternatives considered, actor, timestamp, later changes and final review state. The record should let another person reconstruct the result.

Include one-to-many, many-to-one and many-to-many examples with known answers. Check that the product shows every component, proves that the group totals agree and sends ambiguous groups to review rather than choosing silently.

No. Visibility shows balances and movements across sources. Reconciliation links individual bank movements to accounting records, explains unmatched differences and preserves evidence that the account was reviewed for a stated period.

Our current public pages describe Xero and NetSuite data flows, records kept current and accounting data ready to reconcile. They do not name a standalone AI cash reconciliation agent or publish the evaluation evidence in this guide. Confirm exact integration behaviour and plan availability with us for your setup.

More yield. Less admin. Let us show you.

Skip the months-long implementation. Round is built to plug into your existing stack and start working for you immediately.
Disclaimers:
Nothing on this site is a recommendation to invest. Round does not offer financial advice. If you are unsure about investing we encourage you to speak to a financial advisor. Your capital is at risk when investing. More information here.
Round Financial Limited is authorised and regulated by the Financial Conduct Authority (FRN: 1050315), registered in England and Wales with company number 14609702. Registered office Senna Building, Gorsuch Place, London, E2 8JF, United Kingdom.
Round Financial Limited is an agent of Plaid Financial Limited, an authorised payment institution regulated by the Financial Conduct Authority under the Payment Services Regulations 2017 (Firm Reference Number: 804718). Plaid provides you with regulated account information services through Round as its agent.
Round acts as an Introducer to Insignis Asset Management Limited (Insignis Cash). Round receives a revenue share in return for introducing clients to Insignis Cash. Insignis Cash is a trading name of Insignis Asset Management Limited (Company number 09477376). Insignis Asset Management Limited is authorised by the Financial Conduct Authority under the Payment Service Regulations 2017 (813442) for the provision of payment services.
Keel Money Ltd. Ltd is an Electronic Money Institution authorised by the Financial Conduct Authority under the Electronic Money Regulations 2011 (FRN 1020783). Client funds are safeguarded in UK- or EEA-authorised credit institutions but are not protected by the Financial Services Compensation Scheme. Round Financial Limited is appointed under Regulation 33 of the EMRs to distribute and/or redeem electronic money on behalf of Keel Money Ltd. and is not itself authorised to issue electronic money or provide payment services. More details can be found in the Keel End-User T&Cs, which you must agree to before using any services provided by Keel Money Ltd..
* Rates quoted are the net daily yield from BlackRock ICS Sterling Liquidity Fund as of 13 November 2025. Performance shown as Annual Equivalent Rate (AER) — the annualised rate of return based on daily-compounded NAV growth, including BlackRock fees and Round fees. See pricing page for more details.
** Withdrawal requests must be made by 10:30am for funds to be in your account by the end of the day.
***Assuming your business is eligible for up to £120,000 FSCS protection. Balances over £120,000 per bank will not be protected across all your cash holdings. The Financial Services Compensation Scheme (FSCS) does not cover any e-money products or any products offered by Frost Money Ltd. E-money is not a deposit, savings or investment product and is therefore not protected by the FSCS.
Ratings
G2
4.9 stars
Certificates
Senna Building, Gorsuch Place, E2 8JF