Investors summing partner sales to match a rival’s revenue headline
OpenAI’s August press release claimed an annualised revenue of $40 billion, a figure later corrected to roughly $20 billion after investors were found to have added cloud‑partner sales that the company itself does not count. The episode is not an isolated accounting slip; it is an instance of a recurring practice in which parties reshape a metric’s definition to make a headline number comparable to a rival’s, thereby creating a temporary illusion of parity that fuels valuation talks, media hype, and strategic positioning.
The practice works because two sets of actors—those seeking a favorable comparison and those who aggregate the data for public consumption—operate under different incentive regimes. Investors, whose compensation and reputation depend on the perceived growth trajectory of the firm they back, are motivated to present the company in the best possible light. When a peer such as Anthropic reports revenue that includes cloud‑partner sales, the investors press OpenAI’s finance team to “gross up” OpenAI’s own figure by adding the same class of sales, even though OpenAI’s internal accounting treats those sales as separate. The result is a headline that matches the competitor’s methodology, but the underlying data set is inconsistent. The public‑facing metric therefore misleads analysts, journalists, and other investors who assume a common basis for comparison.
The same dynamic has surfaced repeatedly in other arenas. In the early 2000s Enron used “mark‑to‑market” accounting to record the present value of long‑term contracts as immediate profit, a method that diverged from standard practice and inflated earnings to impress investors. When the discrepancy between Enron’s reported earnings and cash flow became apparent, the company’s stock collapsed and the scandal reshaped accounting standards. The incentive was identical: executives and investors wanted a headline profit line that looked strong relative to peers, and they achieved it by redefining what counted as profit.
A similar pattern appears in the credit‑rating industry that underpinned the 2008 financial crisis. Rating agencies such as Moody’s and Standard & Poor’s assigned “AAA” ratings to mortgage‑backed securities by applying models that assumed low default rates, a modeling choice that differed from the more conservative assumptions used by banks in their internal risk assessments. By publishing ratings based on optimistic assumptions, the agencies produced a metric—credit rating—that could be directly compared with other securities, even though the underlying risk calculations were not aligned. Investors, relying on the ratings, poured capital into the securities, inflating the market for subprime mortgages until the underlying defaults showed that the rating methodology had omitted the risk factors that the banks’ internal models included.
The sports world offers a non‑financial illustration. In Major League Baseball, the “OPS” (on‑base plus slugging) statistic was introduced in the 1970s to combine two separate measures of offensive performance. When teams began to market players based on OPS, analysts retroactively adjusted historical data to compute OPS for eras before the statistic existed, allowing direct comparison across decades. The adjustment required redefining what counted as a “hit” or “walk” in earlier rule sets, effectively inflating the performance of older players to match modern evaluation criteria. The incentive here was to create a headline metric that could be used in contract negotiations and fan discussions, even though the underlying events were recorded under different rules.
Academic publishing also suffers from metric grooming. Journal impact factors are calculated by dividing the number of citations in a given year by the number of citable items published in the preceding two years. In the 1990s, several publishers began to label editorials, letters, and news items as “citable” in the denominator while excluding them from the numerator, thereby lowering the denominator and raising the impact factor. The publishers’ goal was to position their journals as more prestigious relative to competitors, and they succeeded by manipulating the definition of what counted as a “citable” item. When the practice was exposed, the Thomson Reuters database adjusted its methodology, but the temporary inflation had already influenced tenure decisions and funding allocations.
Political polling provides another clear case. During the 2016 U.S. presidential election, some pollsters weighted their samples to over‑represent likely voters in swing states, a weighting scheme that differed from the standard national weighting used by most outlets. The resulting “state‑level” polls displayed a tighter race than the national averages suggested, creating a narrative of competitiveness that attracted media attention and donor contributions. The pollsters’ incentive was to produce a headline that aligned with the expectations of campaign consultants and newsrooms, even though the underlying sample composition was not comparable to other polls.
Even biological research can be drawn into the pattern. In the early 2000s, pharmaceutical firms conducting phase‑III trials sometimes altered the primary endpoint after interim data suggested the original endpoint would not meet statistical significance. By redefining the endpoint—e.g., switching from “overall survival” to “progression‑free survival”—the firms could claim a positive result that appeared comparable to other trials using the new endpoint. The incentive was to preserve the perception of efficacy for investors and regulators, and the practice led to stricter FDA guidance on endpoint pre‑specification.
Across these domains the common causal chain is: a group that benefits from a favorable comparison (investors, executives, teams, publishers, pollsters, drug developers) identifies a metric used by a rival or by the broader market, then adjusts the definition or scope of its own measurement to match that rival’s methodology, even when internal accounting or data collection treats the component differently. The adjustment is performed without transparent disclosure, so external observers assume a common basis. The resulting headline number diverges from the underlying data, creating an information asymmetry that can mislead decision‑makers. When the discrepancy is later uncovered—through internal audits, regulatory review, or independent analysis—the correction often arrives after the inflated figure has already influenced valuations, contracts, or policy choices.
The OpenAI episode mirrors these precedents in its specifics. According to a source with direct knowledge, OpenAI’s investors attempted to produce a “direct comparison” with Anthropic’s annualised revenues. Anthropic’s figure includes sales via cloud partners such as AWS and Google Cloud; OpenAI’s internal reporting excludes those partner sales, treating them as separate revenue streams. By adding the partner sales to OpenAI’s headline, the investors “grossed up” the number to $40 billion, a figure that appeared in August press releases and investor decks. The company later clarified that its actual annualised revenue was $20 billion, roughly $20 billion less than the inflated claim, and that growth remained “more than 70 %” year‑over‑year, a statement that did not compensate for the earlier misstatement.
The technical details of the adjustment are straightforward. OpenAI’s finance team tracks two columns: (1) direct subscription revenue from API customers, and (2) partner‑channel revenue recognized under a revenue‑share agreement with cloud providers. The investors instructed the team to sum column (1) and column (2) for the purpose of the public announcement, despite internal policy that column (2) be reported separately because the cloud providers retain a substantial portion of the gross amount. The public‑facing metric therefore combined two streams that the company’s own financial statements treat as distinct, breaking the alignment between internal reporting and external communication.
The consequences echo those of the earlier cases. Analysts who relied on the $40 billion figure adjusted their valuation models, leading to higher market expectations and a surge in OpenAI‑related venture activity. Media outlets reproduced the headline without noting the methodological difference, amplifying the perception of OpenAI’s market dominance. When the correction emerged, the revised figure forced a reassessment of OpenAI’s competitive position relative to Anthropic and other AI firms, and it raised questions about the reliability of investor‑provided metrics in a rapidly evolving industry.
The persistence of this mechanism suggests that any environment where performance is publicly benchmarked against peers is vulnerable. The incentive to appear on par—or ahead—of competitors can outweigh the risk of later correction, especially when the correction is expected to be absorbed by the market’s attention span. The practice also exploits the fact that most observers lack the granularity to verify each component of a composite metric; they accept the headline at face value, assuming a shared definition.
One might argue that transparent footnotes could mitigate the problem. In practice, the footnotes are often buried in lengthy filings, omitted from press releases, or couched in jargon that discourages casual scrutiny. The Enron case showed that even when detailed disclosures exist, the sheer complexity of the accounting methods can obscure the true economic picture. The rating‑agency case demonstrated that methodological differences can be hidden behind a single letter grade, and the sports‑statistic case illustrated how retroactive adjustments can be presented as “standardized” without explaining the reconstruction process.
The OpenAI incident therefore does not represent a novel ethical lapse; it is a re‑expression of a well‑documented pattern in which actors reshape measurement definitions to engineer a headline that serves their strategic goals. The pattern endures because the cost of adjusting the metric is low—often a spreadsheet change or a revised press release—while the benefit—a higher valuation, more media coverage, or a more persuasive pitch—can be substantial. The risk, meanwhile, is deferred: it materialises only when the inflated number is audited, compared, or otherwise challenged, at which point the correction may already have altered market dynamics.
The broader implication is that any metric used for cross‑entity comparison must be accompanied by a clear, immutable definition that is enforced across all reporting parties. Without such a standard, the same “gross‑up” logic can be applied repeatedly, each time generating a temporary illusion of parity that reshapes expectations and decisions. The OpenAI correction, the Enron earnings inflation, the credit‑rating optimism, the baseball OPS retrofits, the impact‑factor adjustments, the swing‑state poll weighting, and the clinical‑trial endpoint swaps all reveal that the underlying causal process—actors redefining a metric to align with a rival’s figure—operates independently of the specific field, era, or technology.
The final observable fact remains: investors still have a powerful motive to present a revenue figure that matches a competitor’s headline, and the mechanisms for doing so are technically simple. As long as the public relies on headline numbers without scrutinizing the underlying composition, the cycle will repeat. The OpenAI case adds a fresh data point to a lineage that stretches from 19th‑century railroad accounting tricks to 21st‑century AI financing, confirming that the practice is not an anomaly but a durable feature of competitive measurement.