Insights · No. 10

In the AI Free-for-All for Data, What is Proprietary Data?

Proprietary, decision-capturing data over a long period of time is the training signal competitors cannot replicate

"Proprietary data" is the most abused phrase in AI business narratives. Almost every company claims they have it when very few really do. The frontier models, especially OpenAI, are accused by the New York Times and others of stealing what might have been considered proprietary data a few years ago and set a low bar for integrity and stewardship of data sources all in their higher pursuit of AI training progress (NYT). The gap between the claims on what still constitutes proprietary data and the reality is one of the most useful filters there is for distinguishing a defensible AI-native business from a wrapper that has been or will be rapidly commoditized.

Our working definition of proprietary data in the enterprise business context is data that accumulates exclusively through the act of performing the work, that cannot be purchased or licensed from a third party, and that has historical context and increases in value over time. By that standard, a customer list is not proprietary data. A subscription to an industry dataset is definitely not proprietary data. Most of what gets the label fails on all three counts.

Defining proprietary data: value concentrates in decision-capturing, time-dense data (Lateral).

Defining proprietary data: value concentrates in decision-capturing, time-dense data (Lateral).

Three tests

Transaction-generated, not acquired. Genuinely proprietary data is a byproduct of doing the work. Every claim investigated, every load tendered, every contract negotiated produces data that would not exist if the work hadn't been done. As an illustration, a freight brokerage that has moved 500,000 loads over fifteen years holds pricing intelligence, carrier-performance records, and shipper-behavior patterns that exist nowhere else. That data can't be scraped, bought, or synthesized. It was created by the transactions themselves. A company that bought its dataset has something its competitors can also buy.

Decision-capturing, not merely descriptive. The most valuable data records not only what happened but how an expert decided. Descriptive data says a claim was paid; decision data captures why this claim was settled and that one litigated, what the adjuster saw, where judgment overrode the rule. That distinction is everything in an AI context, because the decision trail is what lets a model learn judgment rather than just outcomes. It's also the part a competitor can't reconstruct from the outside. They can see the result, never the reasoning. When a senior broker overrides an algorithm’s carrier recommendation because she knows that carrier struggles with refrigerated loads in July, that override becomes training data. When a subrogation specialist decides to litigate rather than settle because the liability pattern matches previous fraud cases, that judgment is recorded. These decision-capture mechanisms which may have been preserved in workflow systems, case notes, and exception logs constitute the training corpus for AI systems that must eventually exercise similar judgment.

Time-dense, not point-in-time. A single year of data reveals patterns; a decade reveals cycles, regime changes, and the rare edge cases that only appear under specific conditions. An energy-procurement consultant whose data spans the 2008 financial crisis, the 2014 oil-price collapse, the 2021 Texas grid failure, and the 2022 European energy crisis holds training data for scenarios a newer entrant has simply never observed. Time cannot be compressed. A competitor starting today cannot acquire fifteen years of decision-making history regardless of how much capital it raises which is precisely what makes the asset durable.

The 10 categories and where the value concentrates

We have identified 10 distinct types of proprietary data: transaction, workflow, feedback, judgment, outcome, exception, counterparty-intelligence, temporal-pattern, QA, and pricing data. They're not equally valuable. The commodity end is workflow and transaction logs which is useful, but the kind of thing a determined competitor can eventually approximate. The crown jewels sit at the other end: judgment data (how experts decided in ambiguous situations), outcome data (what actually happened over time such as defaults, case results, hire success), and exception data (the edge cases that required a human to step in). Those three are the hardest to accumulate and the most directly useful for training a model that has to operate in a domain where being confidently wrong is expensive.

Prop data x 10

No synthetic data application in vertical enterprise AI

The 2026 workaround for scarcity of training has been to produce synthetic data for training: if the dataset is the moat, simulate it. For descriptive and workflow data in low-risk/low-value consumer applications, that argument has some force and the economics are compelling. For the vertical categories where high-value business decisions are being made and that we focus on, synthetic data will create unpredictable outcomes and further add hallucination, errors and unreliability. The resolution of proprietary data reduces these errors; synthetic data increases the errors.

Frontier models cannot synthesize how a specific carrier's adjusters weighed liability on genuinely ambiguous claims, because that judgment was contingent on context no simulator has. The outcomes cannot be synthesized which loans defaulted, which cases were lost, which hires worked out because outcomes are facts about the world, not patterns to be generated. And the exceptions cannot be synthesized, because the whole point of an exception is that it didn't fit the pattern the generator would reproduce.

Synthetic data is good at manufacturing more of what you already understand. The valuable data is precisely the record of what surprised the experts. Surprise can't be synthesized in advance.

The closed loop of data value accretion

What turns a data advantage into a widening moat is that these categories compound through a closed-loop learning system. Each transaction enriches the dataset. Each expert decision refines the model's grasp of judgment. Each outcome improves predictive accuracy. Each exception sharpens the boundary of where the AI can be trusted to act alone. The result is a flywheel: more clients generate more data, which trains better AI, which attracts more clients. For the owner of proprietary data, the owner's lead grows rather than erodes which is the opposite of what happens to a business who lacks "proprietary data" or relies on data that can be bought by a new entrant.

This is also the cleanest explanation for why horizontal platforms are structurally disadvantaged on data specifically. A platform serving every industry has enormous data volume and almost no depth in any single vertical. It can tell you how organizations broadly use a CRM; it cannot tell you how a specific carrier's adjusters have settled subrogation claims across two decades and four market cycles. In an era where model quality in a narrow domain is gated by domain-specific data, breadth is not an advantage. Depth is and advantage, and depth only accumulates one way.

Where this goes

Proprietary data is the input but not the actionable insight. The next pillar is the one that turns data into something useful and keeps AI pointed at the right problems: domain expertise which is the next scarce input in an age of intelligence.

Sources

Lateral, "The Services Advantage in an AI-Native World". Illustrative examples drawn from the paper.

Up next

The Expertise Lens — why domain judgment is the scarce input when intelligence itself is cheap.

The Durable Value of Vertical Domain Expertise →

Disclosures

The views expressed in this white paper are the best-faith views of the principals of Lateral. This document is not primary research and should not be treated as such. Any such information regarding market forecasts and/or segmentation does not relate specifically to any investment strategy or offering of Lateral. This document is for informational purposes only and reflects the views of Lateral Investment Management as of the date of publication. It does not constitute investment, legal, or tax advice, nor an offer to sell or a solicitation of an offer to buy any security. Past performance is not indicative of future results.

Never miss an issue

Get new Insights by email.

Double opt-in. Unsubscribe any time.