Skip to main content
Resources/Articles/Clean Data Is the Starting Line
Resources/Articles/Clean Data Is the Starting Line
The Loan Ops Room · Part 01

Clean Data Is the Starting Line

What happens after AI reads the document? A year of sessions with banks and loan-servicing teams tell us that extraction is only the beginning.

The Loan Ops Room — what happens after AI reads the document?

What happens after AI reads the document? A year of sessions with banks and loan-servicing teams tell us that extraction is only the beginning. The harder problem is turning what has been read into a clean, linked and trusted record that operations teams can actually use.

A 300-page facility agreement flows into a list of extracted terms — margin, maturity, covenants, repricing dates — all ticked, ending in a circle reading “Then what?”
Extraction ends with a list of ticked terms. The operational question starts straight after it.

Ten minutes into the session, the model has just finished reading a three-hundred-page facility agreement. Every term is on the screen — margin, maturity, the covenant grid, the repricing dates — pulled in about the time it takes to pour a coffee. There’s the usual pause, and someone says a version of “okay, that’s impressive.”

Then the person who actually runs the desk folds her arms. “Fine,” she says. “It can read. So can my team. What I want to know is what happens on Tuesday, when the rate-fixing notice lands and nobody can tell me which tranche it belongs to.”

That’s the question that actually decides things, and I’ve heard a version of it in almost every room I’ve sat in this past year — syndications desks, agency and servicing teams, and private-credit shops. The rooms change; the question barely does. It quietly concedes the one everybody opens with — “can your AI read our credit agreements?” — because reading, the thing everyone came to see, is becoming commoditised. Many tools do it well now. What happens to the data after it has been read is where the time still goes, where the cost still sits, and where the errors still enter the record.

The demo is the easy part

The same scenario has played out, in different forms, in session after session.

A team has already bought a capable extraction tool. It reads documents well; nobody disputes that. And they are still not meaningfully faster. When I ask why, the answer is always some version of the same thing. A notice arrives — a rate-fixing notice, an interest and fee notice, a prepayment notice — and before anyone can do anything with it, a person must work out which deal it belongs to. Which facility within that deal. Which tranche. Whether the figures on it agree with what was booked, and what to do if they don’t. The document was read in seconds. Matching it to the deal took the rest of the afternoon.

Two columns. “What the AI reads” — a facility agreement with margin, maturity, covenants and repricing dates all ticked. Versus “What operations still has to do” — a chain each notice triggers: which deal, which facility, which tranche, does it match what’s booked, what happens next.
Seconds to read the document. The rest of the afternoon to work out where it belongs.

I’ve started asking one follow-up in these sessions, and it lands the same way almost every time. I point at a rate-fixing notice on the screen and ask: “How do you know this belongs to that tranche?” There’s a pause, and then someone — usually the most experienced person in the room — says, “We just know. Someone just knows.” The knowledge is real. It’s also sitting in one person’s head, unwritten, unscaled, and one resignation away from walking out the door.

“Someone just knows”

That’s the moment the conversation actually starts, because it reframes the problem. A credit agreement is not the end of a deal; it’s the first document in a stream that runs for the life of the facility — drawdowns, rate resets, interest and fee notices, amendments, transfers, waivers. Each one arrives separately, often by email, worded differently from the last, and each means almost nothing on its own. A rate-fixing notice in isolation is a page of numbers. Tied to the right facility, the right tranche, and the right repricing schedule, it becomes an instruction the operation can act on with confidence.

This is where extraction quietly stops being enough. You can pull every term from every document and still be holding a pile of disconnected facts. The reading got faster; the connecting didn’t move, because connecting was never a reading problem in the first place. It’s a data problem.

So, the more useful question isn’t “can AI read this?” It’s “can AI connect what it has read to everything else we already know about the deal?”

That connection is the real starting line.

“Information exists in many places” — agreements, emails and inboxes, loan systems — feeding a “linked deal record”: a deal containing Facility A and Facility B, breaking into tranches A1, A2 and B1, with rate-fixing notices, drawdown notices, amendments and waivers hung off each tranche.
A loan as a structured record, not a folder of PDFs — every later notice hung off the right node.

In practice it means treating a loan not as a folder of PDFs but as a structured record: a deal that contains facilities, facilities that break into tranches with their own rates and repricing dates, and every subsequent notice and amendment hung off the right node in that structure. Most of the information already exists somewhere — in the loan system, in someone’s inbox, in the agreement itself. What’s usually missing is the thread that ties it together, so that a machine, and then a person, can follow it from the agreement all the way through to the notice that landed this morning.

Can I trust this number?

There’s a second thing these sessions taught me, and it’s the one the risk and audit people in the room care about most. It isn’t enough for the data to be connected; it must be trustworthy from the first field.

In an operational record, a number that is “probably” right is worse than no number at all, because someone will act on it. So, every extracted term must be traceable back to the exact page, paragraph and supporting language in the source document, and be checked against that source before it ever reaches a person. Not raw model output that happens to read well — but validated data, with its provenance visible behind each field.

A margin field showing 175 bps marked validated, its source listed as Facility Agreement.pdf, page 87, clause 5.2, alongside the highlighted clause it was taken from — and a checklist: connected, validated, traceable to source, ready to use.
Every field traceable to the page, clause and language it came from — and checked against that source.

That is the difference between something that demos beautifully and something you would actually let touch a live book. I’ve watched the mood in a room change the moment people realise they can click a figure and see the clause it came from. That’s when “impressive” turns into “we could rely on this.”

Get it right, and everything downstream gets easier

This is the foundation for much of what comes next. Clean, linked, validated loan data has value well beyond the first document.

Two lifecycle rows. “Get it right”: agreement, notices, amendments, reporting and audit each inherit a trusted record. “Stop at extraction”: every stage becomes reconciliation and the mess moves downstream.
Every later stage inherits either a record it can trust, or the reconciliation you skipped.

Get it right and the benefit compounds across the whole lifecycle — every later stage inherits a record it can trust. Get it wrong, or stop at extraction, and every later stage inherits the mess instead. You end up spending the savings from faster reading on more reconciliation.

This is also, for what it’s worth, where our own work on Smartflow is focused: building a document-derived golden record for the deal, with fields that trace back to its source, reconciled against the bank’s existing records, and delivered into platforms teams already use.

But the principle is broader than our version. Whatever tool you’re looking at, judge it not just on what it can read, but on whether it can connect, validate and operationalise what it has read.

Because reading the document was never the hard part. Connecting it, validating it, and being able to prove it is where the work actually begins.

↑