The Numbers in a Private Equity Deal Get Tested Systematically. The Reasoning Never Has.
Until now.
Tools for testing reasoning have existed for millennia, but private equity has only ever applied that discipline to numbers, never to the arguments underneath them. Large language models are the first thing capable of running that same rigor on reasoning itself, and the shift opens doors the industry didn't know were there.
The Sentence Nobody Checks
Private equity firms buy companies, restructure them, and sell them years later for a profit, betting that today's price is wrong and tomorrow's will prove it. Every bet rests on a model, not in the AI sense, but the older one: a spreadsheet that projects how the business performs over the life of the investment, built from assumptions about revenue, margins, and what the company can be sold for on the way out. Before a deal closes, that model gets torn apart. Analysts run it under a dozen scenarios, question every input, argue over whether the multiple should be 8.5x or 8.7x. The sentence sitting above the model, the one that actually justifies it, "this customer base is sticky," rarely gets the same treatment. It gets a nod around a table and a vote. Nobody asks what's holding that sentence up, because until recently there was no way to check it the way you check a spreadsheet. The model has always depended on the argument, not the other way around, and firms have spent decades scrutinizing the part they knew how to scrutinize while waving the rest through.
Two Categories of Risk, One Left Unexamined
Every fund maintains a process for model risk. Formulas are checked, assumptions are traced to the correct cells, someone runs the scenario where EBITDA comes in 15% below plan. This discipline exists because a model is precise enough to be wrong in ways that can be located and named.
An argument does not offer the same precision. A partner can sense that a memo's reasoning is thin without being able to demonstrate it, and there has been no way to test that reasoning systematically, across every deal, without depending on someone's judgment catching it in a live meeting.
This is logic risk: the risk that the reasoning connecting a firm's diligence to its recommendation does not actually hold. It belongs to the same category as a broken formula. It simply lives in prose rather than a spreadsheet, and until now, no comparable process existed to test it.
The Origin of This Kind of Reasoning
Remove the deal-specific language from any investment thesis and what remains is a form Aristotle would have recognized:
Every business with a captive customer base grows earnings through a downturn. This business has a captive customer base. So it will grow earnings through the downturn.
This is a syllogism: a general rule, a claim that a specific case satisfies the rule, and a conclusion that follows if both premises are true. Stripped of its financial vocabulary, that is what a thesis paragraph is.
In 1847, George Boole formalized this kind of reasoning as algebra, AND, OR, IF, NOT, and that formalization is why a modern trading system can evaluate a logical rule in microseconds. But it also fixed a boundary on what could be automated: language could only be tested once it had been translated into that notation by hand. Numbers acquired a testing infrastructure because numbers were already in a form a machine could operate on. Prose did not, because no system built for algebra could take a paragraph as input.
That boundary has moved. A language model does not require the paragraph to be translated first. It can read "captive customer base" and "grows through downturns" as written, hold the two claims against each other, and determine whether the second follows from the first, or whether it has simply been asserted without support.
What This Kind of Review Looks Like in Practice
Consider a memo built around a healthcare staffing platform. The thesis: recurring placements with hospital systems produce revenue that holds up through a downturn. Read for its logic, the memo raises a specific set of questions:
- The recurring-revenue claim depends on contract renewal rates. Are those rates documented, or is the word "recurring" substituting for a number that was never provided?
- Margin expansion in the model is attributed to operating-expense leverage. Does the diligence document an actual cost initiative, or was the expansion assumed because the model required it?
- One health system accounts for a third of placements. The memo describes the customer base as diversified. Diversified relative to what benchmark?
- The exit multiple matches the entry multiple. Because the sector is expected to re-rate over the hold period, or because no alternative figure was defended?
- The downside case still assumes renewal rates hold at current levels. If renewal rates are the variable under stress, what is the downside case actually testing?
None of this requires re-underwriting the transaction. It requires reading the argument with the same discipline applied to a formula: identifying every point where a conclusion depends on something the document never established, and naming it.
What a System Designed for This Does
Keystone applies this kind of reading to every memo it processes, not only the deals where a reviewer already suspected a problem. It reconstructs the argument: the general rule the memo relies on, the specific claims intended to satisfy that rule, and the conclusion drawn from both. It then evaluates each connection individually. Where a claim is supported by evidence in the document, it identifies that support. Where a claim is asserted without corresponding evidence, it identifies that instead, and points to the precise location of the gap.
The system operates in an isolated environment. No client document is retained, and none is used to train any other model.
The output is not a verdict. It is an inventory of points at which the argument depends on something the memo never demonstrated. A partner reviewing one of these findings may conclude that the underlying assumption is sound and did not require explicit support, that the renewal rate in question is common knowledge to everyone at the table. That is a legitimate conclusion. But it is now a decision made deliberately, rather than an omission discovered only after the business misses guidance well into the hold period.
That distinction, between risk a firm has knowingly accepted and risk no one identified, is close to the entire value of the exercise. It is also worth stating plainly that a system of this kind will not be correct in every instance. Some findings will identify a genuine gap; others will surface a judgment call the deal team had already made correctly. The output functions less as a ruling than as a second reader with no stake in the transaction, and a second reader who is occasionally wrong remains a useful addition to the process.
The Standard This Is Approaching
A fund running its portfolio models by hand today, without spreadsheets or a risk process, would be considered negligent. The tools to do otherwise have existed for decades, and declining to use them is no longer a neutral choice.
The same standard is beginning to apply to the argument beneath the model. It has not yet arrived at most firms, but the condition that prevented it, the inability to test language with the rigor applied to formulas, no longer holds. The firms that close that distance first will not only catch more errors. They will develop the discipline of identifying a flawed argument before it becomes the reason a deal fails, a discipline that compounds the way every other underwriting practice compounds: incrementally, deal after deal, until it becomes the difference between a fund that is surprised and one that is not.