Providers conceal how latency scales with input size, forcing users to run trial‑and‑error tests
The public beta of the Decisions API was tested against two competing back‑ends—Jev, a low‑cost model that had displaced earlier mt0 attempts, and Mercury Decide, a dLLM‑focused service—using fewer than six hundred calls through OpenRouter. The test showed that Mercury Decide’s latency clustered around a 346 ms median and an 860 ms 95th percentile, with occasional spikes that had previously reached 1.3 s but now hovered near 800 ms; Jev consistently responded faster. The pattern that emerged was not a simple matter of one provider being slower, but a failure of the API to scale its response time proportionally to the size of the input, a symptom of a deeper systemic flaw: when a shared compute service is exposed without transparent performance characteristics, users are forced into a black‑box trial loop that distorts provider competition and misallocates resources.
The core of the loop is a mismatch between the information that providers can observe and the information that downstream users require. The Decisions API, like many newly opened services, publishes only an endpoint and a promise of functionality. It does not expose internal queuing policies, model loading times, or the cost of per‑token computation. Users, therefore, must infer performance from empirical measurements. In the test, the user observed that Mercury Decide’s latency did not grow linearly with input length, contradicting the naïve expectation that each additional token incurs a fixed processing cost. The p50 of 346 ms suggests that small payloads are handled efficiently, while the p95 of 860 ms indicates that larger or more complex payloads encounter hidden bottlenecks—perhaps model warm‑up, dynamic batching, or variable backend scaling. Because the API offers no telemetry, the user cannot diagnose whether the slowdown stems from internal resource contention, from the provider’s choice to prioritize certain request classes, or from the underlying hardware’s saturation.
When such opacity exists, two forces converge. First, providers have an incentive to hide performance details that might reveal inefficiencies or competitive disadvantages. By presenting a single latency figure or a vague “fast enough” claim, they hide how latency grows with input size, preventing users from comparing providers on actual scaling. Second, users, lacking reliable metrics, resort to brute‑force experimentation: they send repeated calls, compare median and tail latencies, and adjust request size or provider choice based on observed outcomes. This iterative probing consumes compute cycles, network bandwidth, and developer time—resources that could otherwise be allocated to productive work. Moreover, the loop creates a feedback distortion: providers that happen to perform well under the limited test conditions receive more traffic, reinforcing their market position, while those whose performance degrades only on larger inputs remain invisible, even if they could be more cost‑effective at scale.
The black‑box latency loop is not unique to modern AI APIs. In the early 1980s, the ARPANET community faced a similar opacity. The network’s packet‑switching hardware offered no public metrics on congestion or round‑trip time, leading application developers to experience unpredictable delays. The resulting frustration spurred the creation of the “ping” utility in 1983, a simple program that measured round‑trip latency by sending echo requests. Ping’s adoption revealed that many network paths exhibited latency spikes that were invisible to higher‑level protocols. The lack of transparent performance data had forced users into a trial‑and‑error regime that wasted bandwidth and hampered early distributed applications. The eventual standardization of latency measurement tools broke the loop by supplying the missing information, allowing both network operators and application developers to optimize routing and load balancing.
A parallel episode unfolded in telecommunications during the 1920s. The American Telephone and Telegraph (AT&T) long‑distance network charged rates based on a complex, internally calculated cost model that accounted for line attenuation, relay usage, and regional traffic patterns. The calculation method was not disclosed to corporate customers, who therefore could not predict the expense of routing high‑volume calls across the network. Companies responded by establishing private “carrier” agreements with intermediate operators, attempting to infer the hidden cost structure through trial calls and billing statements. This practice introduced inefficiencies: traffic was rerouted through suboptimal paths, and the market for inter‑carrier arbitration grew to manage disputes. Only after the Federal Communications Commission mandated rate transparency in the 1930s did businesses gain the ability to align routing decisions with actual cost, eliminating the costly black‑box probing phase.
Finance offers a more recent illustration. The London Interbank Offered Rate (LIBOR) was derived from a panel of banks that submitted daily estimates of borrowing costs. The methodology for aggregating these submissions was opaque, and the banks had a direct incentive to influence the rate to benefit their own positions. Market participants, unable to verify the underlying calculations, relied on historical patterns and occasional arbitrage opportunities to gauge the reliability of LIBOR. The resulting uncertainty contributed to mispricing of derivatives and prompted regulators to demand greater transparency, ultimately leading to the replacement of LIBOR with risk‑free rates that are observable from market data.
In cloud computing, Amazon Web Services introduced EC2 spot instances in 2009. Spot instances allow users to bid on unused capacity at discounted rates, but the pricing algorithm and the likelihood of interruption are not disclosed in detail. Users must monitor spot price histories and design fault‑tolerant architectures to survive unexpected termination. The lack of explicit performance guarantees forces developers to embed continual health checks, checkpointing, and redundant workloads—overheads that negate some of the cost savings. The pattern mirrors the black‑box latency loop: a service that promises cheap access without revealing its operational dynamics compels users to build complex mitigation layers.
The black‑box latency loop also appears in biological regulation. In the human endocrine system, hormone release often follows feedback loops that are not directly observable without invasive testing. Physicians, lacking real‑time hormone concentration data, prescribe medication based on periodic blood tests, adjusting dosages iteratively. The hidden dynamics of hormone synthesis and receptor sensitivity create a medical analogue of the computational latency loop: treatment decisions are made under uncertainty, leading to cycles of dosage adjustment that consume clinical resources and risk suboptimal outcomes.
Across these domains, the common causal chain is: a provider exposes a service behind an interface; the provider withholds detailed performance or cost metrics; downstream users, needing those metrics to make efficient choices, engage in empirical probing; the probing consumes resources; providers that happen to perform well under limited probes capture disproportionate demand; the market fails to allocate resources based on true efficiency. The loop persists because the provider’s incentive to protect proprietary details outweighs the user’s desire for transparency, and because the cost of building systematic measurement tools often falls on the user rather than the provider.
The decision‑evaluation test that compared Jev and Mercury Decide embodies this loop. The user’s measurement of 346 ms median and 860 ms 95th percentile for Mercury Decide, together with the observation that latency does not increase linearly with input size, provides only a snapshot. Without access to Mercury Decide’s internal scaling policy—whether it employs dynamic batching, model sharding, or tiered hardware—the user cannot predict how the service will behave for larger workloads or under sustained traffic. Consequently, the user must continue to issue test calls, varying input length, token composition, and concurrency, to map the hidden performance surface. Each test consumes compute cycles on the MacBook Neo, occupies network bandwidth, and adds latency to the user’s own workflow. Meanwhile, Jev’s lower latency, achieved through a simpler, cheaper model, may be selected not because it is intrinsically superior for the intended task, but because its performance is more predictable given the limited data.
The loop’s endurance is reinforced by the public‑beta status of the Decisions API. Beta releases often lack service‑level agreements (SLAs) and are exempt from strict reliability guarantees. Providers can therefore adjust internal resource allocation without notifying users, further widening the information gap. The user’s note that Mercury Decide “scaled far more consistently with size” after a recent reduction of extreme latencies suggests that the provider altered its backend—perhaps by adding caching or adjusting batch sizes—without publishing a changelog. The user’s only recourse is to observe the new latency distribution and infer the change, restarting the trial cycle.
If the loop were broken, the user would receive real‑time performance telemetry directly from the API: per‑request queue depth, estimated processing time, and model load factor. Armed with this data, the user could compute an expected latency function \(L(s) = a + b \times s\), where \(s\) denotes input size, and compare it across providers without exhaustive probing. Providers would be compelled to improve their latency curves to retain traffic, and the market would allocate workloads based on measurable efficiency rather than opaque reputation.
The persistence of the black‑box latency loop raises a final, unresolved fact: despite the proliferation of open‑source monitoring tools and community‑driven benchmark suites, major platform providers continue to withhold fine‑grained latency metrics for their flagship APIs. Whether this omission is driven by competitive secrecy, engineering constraints, or a belief that abstract performance guarantees suffice remains opaque. The loop therefore endures, compelling each new generation of developers to repeat the same costly probing process that has recurred from ARPANET to spot instances, from LIBOR to endocrine therapy.