Large Language Thing

Home/Concepts/Metacognition in venture capital

Metacognition in venture capital

Second-order knowledge is parasitic on first-order flow. To know that a belief has become unreliable, something must have reached you that the belief did not predict. No amount of…

The child who could not know

John Flavell was watching young children fail at a task that had nothing to do with intelligence. In the early 1970s he asked them to study a set of items until they were confident they had memorised them, and then to say so. Many stopped early, certain, and were wrong. It was not that they had failed to learn the items. It was that they had failed to know whether they had learned the items. Flavell called this gap metacognition — cognition about cognition — and split it into two moving parts: monitoring, the judgement of one's own state, and control, the action taken on that judgement. Thomas Nelson and Louis Narens later gave the split a formal architecture in 1990: monitoring flows upward from process to a model of the process, control flows downward from that model back into the process, and the two form a loop. A child who cannot monitor cannot know when to keep studying. The failure is structural, not a deficit of effort.

The detail that matters for everything downstream is this: monitoring is not a fact stored about oneself. It has to be a reading, taken now, from a signal that is still arriving. A child who memorised a list yesterday and has received no further contact with it has no way to know today whether it has decayed, except by testing again. Knowledge of ignorance has an expiry date set by the last time the signal was checked.

The thesis that outlived its market

Venture investing has its own version of the child with the list. A partner builds a thesis — say, that a certain category of infrastructure tooling will consolidate around three winners because of network effects in a specific developer workflow. The thesis is built on filings that show capital concentrating, hiring signals that show senior engineers moving toward two of the three candidates, product telemetry showing usage curves bending upward, and a market structure that, at the moment of underwriting, has a small number of credible entrants and a large greenfield of undecided customers.

Eighteen months later the market the thesis assumed has dissolved. A platform shift made the workflow itself obsolete. A regulatory change altered the unit economics for the whole category. A larger incumbent absorbed the customer base through a bundling move nobody priced in. None of this is exotic; it is close to the median outcome for a venture thesis with a multi-year horizon. What is exotic, and what recurs with depressing regularity, is that the thesis continues to be defended for a year after the dissolution — reflected in follow-on decisions, board votes, and reserve allocations — because nothing forced a re-reading. The partner's confidence in the thesis was formed once, at underwriting, from a snapshot of filings and signals current at that time. It was never designed to be revisited unless something dramatic and undeniable forced the question. Ambiguous decline does not force the question. It just accumulates.

This is Flavell's child, scaled to a fund's balance sheet. The partner is not lacking analytical skill. The firm's diligence at entry was often excellent — a defensible, well-reasoned read of the data available at the time. What is missing is a live monitoring channel running after the decision, something that would generate a dated signal saying: the thing you believed at month zero no longer matches what is arriving at month eighteen. Without that channel, confidence is a report about the past dressed as a report about the present.

Where the intake actually is

Venture firms are not short of data. They stream filings — Form D, cap table changes, subsequent financing rounds. They stream hiring signals — job postings, LinkedIn movement, executive departures. They stream product telemetry when portfolio companies share dashboards, and often that telemetry is contractually available in near real time. They stream market structure — competitor fundraising, category consolidation, customer churn reported informally through the network. All four streams exist and, in well-run firms, are collected.

The failure is not collection. It is that the streams are consulted episodically — at board meetings, at follow-on decisions, at annual portfolio reviews — rather than treated as a continuous input against which the thesis is standing prediction error. A board meeting is a scene, bounded in time, much like the episode in which a Large World Model receives live sensing: for the duration of the meeting, the partner's model of the company can be checked against fresh numbers, and error can be detected. But the meeting ends. The knowledge gained about where the thesis was wrong does not persist as an active monitor; it becomes a memo, filed, and the next check is months away. The gap between checks is exactly where a market can dissolve unnoticed.

The lineage, seen from the term sheet

This is the same structural ladder the whole series has been climbing, and venture makes its rungs unusually legible because the cost of a missed rung is denominated in dollars and dated in quarters.

A model built once from a static dataset — comparable to a Large Language Model's frozen corpus — can be calibrated in aggregate. A firm can look back over a hundred past investments and show that its stated confidence levels tracked realised outcomes reasonably well across the portfolio. That is a real form of calibration, and it is worth having. But it says nothing about any single live thesis, because the corpus that produced it contains no news about what is happening to that thesis now. The firm's back-tested win rate does not update when a competitor raises a surprise round on a Tuesday.

A firm that runs quarterly deep-dives — pulling telemetry, re-underwriting the market map, stress-testing the thesis against the current filings — has a genuine monitoring channel, comparable to a Large World Model's prediction error against live sensing. For the duration of the review, the firm can detect that its model of the company is failing. But the channel closes at the end of the quarter. Whatever was learned about the thesis's blind spots does not stay active; it has to be rediscovered next quarter, from scratch, if anyone remembers to look.

The position that closes the gap is the one where the streams never stop and the beliefs carry provenance: this filing changed the revenue assumption; this hiring signal changed the competitive-moat assumption; this telemetry drop changed the retention assumption; each attributed, dated, and revisable the moment a contradicting signal arrives, rather than at the next scheduled review. That is the Large Universe Model position, applied to a portfolio: confidence in a thesis is not a number set at underwriting and revisited by calendar. It is a maintained quantity, continuously checked against everything still arriving, with a record of which source moved it and when. Knowing that a thesis has gone stale becomes a subscription, not a one-time diligence purchase.

Objections a partner would actually raise

Our models are already calibrated. We back-test every thesis category against realised outcomes and our hit rates are within a few points of stated confidence. That is metacognition; you are asking for something statistics has already delivered.

This is true and worth taking seriously. Back-tested calibration on a stable category is a genuine achievement. It holds under one condition: that the market a new thesis enters resembles, in relevant structure, the markets in the back-test sample. That condition is exactly what a static back-test cannot check, because checking it requires fresh evidence about the market the thesis is currently in — the very stream that stops the moment diligence ends. A firm's aggregate calibration can be excellent while every individual thesis inside it silently exits the distribution the calibration assumed. The statistics do not remove the need for an ongoing signal; they relocate the risk to an assumption that only continued observation can audit.

Investors are famously overconfident already — deal fever, sunk cost, pattern-matching on the last win. More continuous data will just give overconfidence more material to rationalise with.

Granted, and the record bears this out: partners anchored on a strong founding team tend to keep faith with a thesis well past the point the market data would justify, because conviction is read from fluency of the pitch rather than from the telemetry. That is not an argument against continuous monitoring; it is a description of monitoring substituting an internal proxy — conviction, pattern-match, story coherence — for an external outcome signal. Where the outcome signal is forced to arrive on a schedule regardless of story — mark-to-market pricing in liquid markets, for instance — professional forecasters are well calibrated. The lesson is that monitoring is only as good as the feedback forced through it, which is an argument for building the feedback channel properly, not for abandoning the idea that one is needed.

A quarterly memo is not a monitor; it is a photograph, and photographs do not notice when the room changes after the shutter closes.

The harder objection is the third: that continuous streams bring noise, correlated errors and manipulated signals — a founder gaming telemetry, a hiring spike engineered for a raise — and that a firm chasing every update may end up less reliable than one anchored to disciplined periodic review. This is fair, and it is why provenance, not volume, is the load-bearing part of the argument. A belief about a thesis that records which filing, which telemetry feed, which hiring signal moved it, and when, can be checked and discounted if the source turns out corrupted. A firm that reacts to undifferentiated noise without that attribution is worse off than one that waits for quarterly clarity. Continuous intake is not an improvement by itself. Continuous, attributed intake is the only arrangement in which a partner can tell, after the fact, whether the market moved or the microphone did.

Continue