Guide

The model was wrong about the covenant. Here is why.

An analyst pastes a credit agreement into the firm’s AI tool and asks whether the leverage covenant is breached. The answer is fluent, confident and wrong, because the agreement defines EBITDA with six add-backs the model never saw. The analyst catches it, and stops trusting the tool on anything they could not check themselves.

That failure has five causes, and all five are fixable by the person asking. This guide goes through them in the order we teach them in a workshop, on a firm’s own documents. None of the fixes requires a different model, and one of them requires the model to say less.

Where the errors come from

why is chatgpt inaccurate on financial documents

It answered from memory, because nobody gave it the document.

A model asked about a covenant with nothing in front of it answers from what covenants usually say. Most of the time that is close enough to be dangerous. In investment work the value sits exactly in the places where this deal departs from the market: the negotiated definition, the carve-out, the basket, the date that is one quarter later than usual. A model answering from the general case misses every one of them, and it does so in the same confident register it uses when it is right.

The same thing happens when the document is given but the question is not. “Summarise this CIM” produces a summary of a CIM. It does not produce the analysis this firm runs on a CIM, because the firm’s process is not in the prompt. The output reads as inaccurate when it is really answering a different question from the one the analyst had in their head.

The five fixes

how to improve ai accuracy on contracts and deal documents

Five things, in the order they pay.

Each of these cuts a class of error on its own. Together they are the difference between a tool your people stopped using and a first pass a partner will read.

Give it the document, whole
The agreement, the definitions schedule, the amendments, and the compliance certificate, in the tool, in one conversation. A model reading the actual text answers from the actual text. Most accuracy complaints we hear turn out to be a question asked about a document the model was never shown.
State the process, not the task
Tell it how this firm reads this kind of document: what the deal team pulls from a CIM first, which definitions the credit team checks before any ratio, what counts as an exception on this desk. An analyst needs to teach the tool not just to summarise a CIM but to analyse it according to the firm's own process. That instruction is the most valuable prompt the firm owns, and it is usually in somebody's head.
Ask for the citation
Every claim with the clause, page or cell it came from. This does two things. It makes the output checkable in seconds rather than minutes, and it changes the model's behaviour: an answer that has to point at its source is far less likely to be an answer from memory.
Make it say what it could not find
Instruct it that a question the document does not answer gets 'not in the document' rather than a best guess. The most expensive errors are filled gaps. A model that is allowed to say less is more accurate than one that must always produce an answer.
Check it the way a partner would
Read the output against the source for the three things that matter on this desk, every time, before it goes anywhere. The reading becomes a structured first pass across the whole book and the person spends their time on the exceptions. The judgment stays with the human whose name is on the work, which is where it was always going to stay.

Which model

which ai model is most accurate for finance

The one your firm has already approved.

The frontier models are close enough on reading and reasoning that the choice between them matters less than the five fixes above. A well-asked question on a mid-tier model beats a bare question on the best one. What does matter is that the tool is inside the firm’s own data terms, open on the desk where the work happens, and the same tool your people open every morning, because practice is what makes the habits hold. We teach on whatever the firm has licensed. We are an Anthropic partner and the workflows we help clients build run on Claude, but nothing in this guide depends on it.

What it will not fix

when not to use ai in private equity

Judgment stays where it was.

None of this makes a model the person who signs. A covenant check with a citation is a first pass a credit officer reads faster, and the officer still reads it. A triaged inbox puts the right email in front of the right person; it does not reply. The work this helps is the work with reading in it: reading to a decision, assembling from many sources, reconciling one record against another. Where the value of the work is the judgment itself, the model is at most a second reader, and a firm that forgets that has traded one accuracy problem for a worse one.

The four patterns where AI earns its place in an investment firm, and the limits, are on Industries.

Questions

The questions that follow.

What we are asked once a firm accepts that the errors have a cause.

How do I stop AI hallucinating on our deal documents?

Give it the whole document in the conversation, tell it to cite the clause or page for every claim, and instruct it to answer 'not in the document' when the document does not answer. Those three instructions remove most of what gets called hallucination on investment work, because most of it is a model answering from the general case about a document it was never shown.

Is AI accurate enough to read a credit agreement?

As a first pass, with the definitions schedule and the amendments in front of it and a citation on every claim, yes, and the credit officer still reads it. As a substitute for that reading, no. The gain is a structured first pass across the whole book, with the person's week spent on the exceptions.

Does a bigger or newer model fix accuracy?

Less than the question does. A well-asked question on a mid-tier model beats a bare question on the best one. The frontier models are close on reading and reasoning; the gap between firms comes from whether the document, the process and the citation were in the prompt.

How do we make a whole team ask better questions?

Teach them on their own documents, in a session where each person does one piece of real work correctly with somebody watching, then write down the process instruction that worked for that desk. That is what a prompt engineering workshop on this site is. A generic course on generic material does not transfer.

Should we build a retrieval system before anything else?

Not before the people who will use it can get an accurate answer on a single document. Retrieval helps once the questions are good; it finds the right pages for a question that still has to be asked well. Most firms we meet have a retrieval problem later and a question problem now.

How do we know the accuracy improved?

Measure it on your own volume. Take a set of documents the desk has already read, run the first pass with the five fixes, and count the errors against the desk's own reading. That number, before and after, is the one a partner will believe, and it is how we measure a workflow engagement.

Bring us one document.

The one the tool got wrong, and the person who caught it. We will show you, on that document, where the question went and what the first pass looks like when it is asked well.

[email protected]

Subject line: Accuracy