Everett Cedarholm View the notebook

Work 2026 AI / Investing

Exit Probability Model

Given only what was knowable on the day a startup raised its Series A, can you rank which ones go on to exit? On historical Crunchbase data the answer is yes, about 3.5× better than random — and that ranking still is not an investment decision.

Role
Data & modeling
Year
2026
Source
Crunchbase snapshot
Trained on
4,898 Series A cos.
Target
Exit in 5–7 yrs
Best AUC
0.83
Decides for you
No — see below

The
question

A venture fund sees far more Series A companies than it can diligence properly. The screening problem is one of ordering: which of these deserve the week of work, and which can wait.

So the question I set was narrow on purpose. Take a company as it stood on the day of its Series A — round size, prior funding, age, sector, geography, nothing after that date — and estimate the probability it reaches a meaningful exit, an IPO or a real acquisition, within five to seven years.

The reason to build it on historical data is that the answers already exist. Companies that raised an A in 2008 have long since exited, folded, or quietly continued. You can train on what happened and check yourself against it.

That is also the trap. A dataset where the outcomes are already known is a dataset that has been shaped by those outcomes — and most of this project turned out to be about how badly that shaping distorts the answer.

Cleaning the data

54,000 → 4,898

The raw export is a wide table of companies with funding columns, status, location, and category tags. Very little of it was usable as delivered. Each step below throws away rows, and each one is a judgment call that changes the answer.

Read the numbers as numbers

Funding totals arrived as text with inconsistent digit grouping — 17,50,000 is 1.75 million, not seventeen. Parsed as ordinary thousands separators, every large round would have been silently wrong.

54,353

Find the companies that actually raised an A

In this export round_A = 0 means "never raised a Series A", not "raised nothing". Reading it as a real zero would have dropped 40,435 unrelated companies into the training set as failures. Another 4,856 were simply blank.

9,003

Keep only companies with a founding date

Founding year is blank on 29% of rows, and without it there is no way to place a company in a cohort or measure its age at the A. Restricting to those founded 2005–2014 leaves the usable window.

5,973

Train only on cohorts old enough to have an answer

A company founded in 2011 had until roughly 2018 to exit. Anything more recent would be labelled "no exit" purely because not enough time has passed, so training stops at 2011 and the 1,075 companies from 2012–2014 are held back to score.

4,898

Two more decisions did not remove rows but changed what they mean. Defining the label: 713 companies show status acquired, but many acquisitions are talent purchases that return nothing — filtering on signals available at the A, a round of $3M or more plus prior institutional backing, reclassified 189 of them as non-exits. Handling the gaps: state is missing on 44% of rows and region on 19%. Rather than impute, missingness became its own feature — a thin profile is itself information, though as it turns out, information about the wrong thing.

What supervised learning is

60 features, one label

Supervised means you learn from examples where the answer is already attached. You hand the model thousands of rows, each one describing a company at a fixed moment, and each carrying a label saying what eventually happened to it. The model searches for combinations of those numbers that separate the labelled ones from the rest, and you keep the version that works best on rows it was never shown.

Nothing in that process understands startups. It has no idea what a founder is. It is fitting a function to labelled history, on the assumption that the future resembles the past — which is exactly the assumption venture capital exists to violate.

The hard part is rarely the model. It is deciding what the label means, and making sure every input was genuinely knowable before the outcome. A single feature that leaks post-Series-A information will lift the score dramatically and teach you nothing.

Four model families were trained and tuned by cross-validated grid search: logistic regression as an interpretable floor, then random forest, XGBoost, and LightGBM. All of them were class-weighted, because exits run about one in ten and an unweighted model would simply predict "no" every time and be right 90% of the time.

One row of training data Simplified
  • Series A size$6.0M
  • Pre-A capital raised$1.2M
  • Seed-to-A step-up5.0×
  • Days, founding to first cheque612
  • Age at Series A2.4 yrs
  • Sector base exit rate12.4%
  • Tech hub tier1
  • Fields missing on profile2
  • …and 52 more
  • Did it exit within 5–7 years?Yes

Of 5,973 companies, 539 cleared the bar — 9.0%, of which just 15 were IPOs and 524 were acquisitions judged meaningful. That imbalance, roughly 8.5 to 1, shapes every decision downstream.

What it found

980 held-out companies

All four models landed in the same band, and a weighted ensemble of them edged slightly ahead. The headline is real: the model does separate signal from noise, and it does it on a question that is genuinely hard to call by eye. Read every number below with one caveat attached, though — they all measure whether a company exited, and none of them measure what the exit was worth.

Discrimination 0.83

AUC for the ensemble. Given one company that exited and one that didn't, it ranks them correctly 83% of the time. 0.50 is a coin flip.

Lift at the top 3.5×

The top-scored 10% exited at 36.7%, against a 10.5% base rate in the test set. Lift on exits, not on dollars returned.

Still wrong 63%

Share of that same top decile that did not exit. The best bucket the model can find is still mostly misses.

Seed variance ±0.02

AUC swing across ten random seeds, ranging 0.798 to 0.872. Same data, same code, different shuffle.

Actual exit rate, by predicted decile 98 companies per bar
12345 678910
← Lowest predicted probability Highest →
Read the ends, and then read the middle. The bottom two deciles contain no exits at all and the top two are far above base rate, which is genuinely useful — it means the model is good at recognising companies that were never going to make it. But deciles 7 and 8 are inverted, and 3 through 8 are packed into a narrow band. For everything that is not obviously bad or obviously good, the ordering is close to arbitrary.
ModelAUC Precision @ top 10%Lift
Logistic regression0.80735.7%3.40×
Random forest0.81338.8%3.69×
XGBoost0.82435.7%3.40×
LightGBM0.81734.7%3.30×
Weighted ensemble0.82736.7%3.50×

Notice how little separates them. A spread of 0.02 AUC across four very different model families, against a ±0.02 swing from reshuffling the data, says the ceiling is set by the dataset rather than the algorithm.

Why I
wouldn't
decide
off it

The prediction works. That is the part I want to be clear about first: the model finds companies that go on to exit, it does it 3.5× better than chance, and four independent model families agree on the ranking. If the job were "tell me who exits", this would be a solved problem on public data.

The job is not that. An investor is not buying exits, they are buying the size of one — and the moment the target changes from whether to how much, this ranking stops being usable. That is the first fault below, and the four after it are the reasons I would not lean on it even for the narrow question it does answer, roughly in order of how much they matter.

01

An exit is not a return

The model predicts whether a company exits, not what the exit pays. In this label a billion-dollar IPO and a two-engineer acquihire are the same thing: a 1. So is the far more ordinary case — a company acquired for $20M after raising $18M, where the founders get a soft landing, the acquirer gets a team, and the fund gets its money back without the return that justified the risk.

Those outcomes are not rare enough to wave away. They are the modal acquisition. Venture returns come almost entirely from the extreme right tail, and a binary exit flag cannot see the tail — it fires on the acquihire and the outlier with identical confidence.

I cannot correct for it here, and that is the real constraint. This Crunchbase snapshot records that an acquisition happened, not the price it happened at. With no deal values there is no way to weight a success by what it returned, or even to count how many of the 524 acquisitions in the training set closed below the capital that went in. The best available proxy was to filter at the label — a round of $3M or more plus prior institutional backing reclassified 189 acquisitions as non-exits — but that screens on inputs, not on outcome. It is a guess about which acquisitions were probably real, standing in for a number the dataset never contained.

Which means a fund run off this signal would quietly optimise for safe mid-size acquisitions: the outcomes the model scores as wins and the fund books as losses.

No deal values in the data — every exit weighs the same

02

Half the negatives are not really negatives

Crunchbase marks a company operating until somebody bothers to update it. So the "did not exit" class is a blend of companies that genuinely failed and companies that are still perfectly alive and simply have not exited yet. Nobody files paperwork when a startup quietly winds down.

The model is therefore trained to predict something closer to "did this company generate a recorded liquidity event" than "was this a good business" — and it learns the difference.

Known negatives are far outnumbered by unlabelled ones

03

Missing data tracks attention, not quality

Founding year is absent on 29% of rows and state on 44%. That absence is not random. Companies covered by tech press get complete records; companies outside that orbit do not, regardless of how they perform.

Treating missingness as a feature improved the score, which should have been the warning. Part of what the model rewards is having been noticed, and being noticed is downstream of the same coastal, well-networked conditions that already correlate with exits.

04

It degrades the moment time moves

Training on 2005–2009 and testing on 2010–2011 drops AUC to between 0.76 and 0.78, and one model fell far enough to be flagged. The measured exit rate falls from 13.1% to 5.4% across those cohorts — not because the later companies were worse, but because the observation window is shorter.

Deployment is exactly this problem, permanently. You are always scoring the cohort with the least time on the clock.

AUC 0.83 → 0.76 across a five-year shift

05

None of the inputs are what investors underwrite

There is no founder history in this data, no lead investor, no syndicate quality, no product, no revenue, no exact round dates. What remains is round size, timing, sector, and geography — the outside of the box.

In cross-validation, cutting the feature set from 60 down to 20 barely moved the score, and cutting all the way to 5 cost about 0.09 AUC. Most of those 60 columns are restatements of a handful of facts, and the handful is thin.

What survives is narrow but real: the model is reliable at identifying companies unlikely to exit, and the bottom fifth of its ranking contained no exits at all. That makes it a defensible way to order a pipeline — decide what to read first, decide what can wait. It is not a way to decide what to back, because the question it answers and the question a fund asks are not the same question.

How to make it real

In order of leverage

The fixes worth doing are not modelling fixes. Tuning was already exhausted — the useful moves are to the target, the framing, and the inputs.

ChangeWhy it moves the needleCost
Predict return multiple, not exit

Replace the binary label with exit value over capital raised, and the target stops rewarding acquihires. It is the only change that makes the output answer the question a fund is actually asking — and the one change this dataset cannot support, because it has no deal values in it.

Needs deal values
Model time-to-exit, not exit

Survival analysis handles censoring properly, so a 2013 company that has not exited yet stops being treated as a confirmed failure. This is the correct statistical frame for the problem and removes most of fault 04.

Free
Add investor identity

Lead investor and syndicate track record are, plausibly, the strongest single predictor available and are entirely absent here. Also the most confounded — good investors pick well and open doors, and the data cannot separate the two.

New data
Add founder history

Repeat founders, prior exits, and team pedigree at the time of the A. Standard in every human diligence process, missing from every column here.

New data
Exact round dates

The dataset gives first and last funding dates only, so round cadence had to be approximated. Precise dates sharpen every velocity feature.

New data
Cut to 20 features and calibrate

Twenty features cross-validated as well as sixty. A smaller set plus isotonic calibration produces probabilities that can be read as probabilities, which matters more than a further 0.005 of AUC.

Free

None of this is exotic, and that is rather the point. QuantumLight, the fund Nik Storonsky started after Revolut, sources deals this way and closed $250M on the premise in 2025, saying every investment it had made to that point came out of a model rather than a partner's network. The technique is not the moat. What I ran here is four standard classifiers, a grid search, and a public snapshot — the gap between this and a working screen is the data, not the algorithm.

Which is the honest summary of the exercise: a supervised learning problem, taken end to end on data anyone can download, to see how far public information alone gets you. It gets you a real ranking and a target worth arguing with. Most of the work turned out not to be modelling — it was deciding what the label means, proving no input leaked from the future, and then being clear about which of the resulting numbers I would actually stand behind.