q08

The fixed‑size hash that ignored I/O

2026-10-02 · Git 3.0's upcoming SHA-256 default will

The proposal to make SHA‑256 the default hash in Git 3.0 has been called a costly mistake. The mistake stems from optimizing the hash representation for stack residency while disregarding the I/O costs that dominate any real‑world Git operation.

Programmers who championed the change pointed to the concrete advantage of being able to treat a hash as a fixed‑size record that fits comfortably on the CPU stack. In the discussion they noted that such a layout is “incredibly efficient” because it eliminates the need for dynamic allocation and lets the hash be copied with a single move instruction. They also observed that this efficiency dates back to the era when MD5 was the standard, a time when collision attacks on MD5 were already well documented and widely understood. The argument rested on the assumption that the hash is an internal datum whose manipulation cost determines overall performance.

What the argument omitted is the way a Git command actually spends its time. Whether the operation reads a blob from disk, writes a pack file over the network, or even accesses a page‑cached inode, the latency of moving data to or from storage dwarfs the few nanoseconds saved by keeping the hash on the stack. As one commenter put it, “as soon as you read from or write to a disk or a network or even memory, any cost saving is completely gone.” In other words, the local optimization of the hash’s representation is swamped by the systemic cost of I/O, which is the true bottleneck for version‑control workloads. The error is therefore not a flaw in the hash function itself but a mismatch between the metric that guided the change (stack‑resident size) and the metric that actually determines end‑to‑end latency (data movement).

This pattern recurs whenever a designer improves a component by optimizing a narrowly measured attribute while ignoring the attribute that couples the component to the rest of the system. In mechanical engineering, early steam‑engine builders increased boiler pressure to extract more work per cycle, celebrating the gain in thermodynamic efficiency. They paid less attention to the fatigue life of the riveted shells, which operated under cyclic stress far beyond what the material could endure. The result was a series of catastrophic boiler explosions in the mid‑19th century, prompting the introduction of safety valves and regular hydrostatic testing. The engineers had optimized a local variable—pressure—while the system‑level failure mode depended on stress accumulation, a factor invisible to the pressure gauge.

In evolutionary biology, certain fish species evolved extremely high fecundity, laying millions of eggs per season. The trait was favored because it maximized the number of offspring that could reach maturity in a predator‑rich environment. However, the energetic cost of producing so many gametes reduced the mothers’ somatic reserves, making them more vulnerable to starvation during lean years. Populations of these fish therefore exhibited dramatic boom‑bust cycles that correlated poorly with the simple fecundity metric. The evolutionary pressure had optimized a local reproductive output while ignoring the coupling to individual survival, a systemic cost that only manifested when environmental conditions shifted.

In the realm of economics, manufacturers in the late‑20th century pursued labor‑cost reductions by relocating assembly lines to regions with the lowest hourly wages. The decision was justified by the narrow accounting metric of unit labor expense. What the calculation omitted were the hidden coordination costs: longer supply chains, increased inventory buffering, quality‑control overhead, and the erosion of tacit knowledge that resided in the original workforce. When those hidden costs rose—due to political instability, rising freight rates, or intellectual‑property leakage—the apparent advantage vanished, and many firms reversed course, a phenomenon later labeled “reshoring.” The manufacturers had optimized a local cost driver while the system‑level profitability depended on a bundle of factors that were not captured by the wage figure alone.

Legal systems exhibit a similar drift. Legislators sometimes craft statutes with highly precise definitions to reduce judicial interpretation and thereby lower litigation expenses. The precision is celebrated as a way to make the law “more efficient.” Yet the same precision can create loopholes: regulated actors learn to structure their behavior just outside the defined boundaries, forcing courts to expend resources on complex factual inquiries to determine whether a technical violation occurred. The net effect can be an increase, not a decrease, in the administrative burden of enforcement. The lawmakers optimized a local metric—textual specificity—while the systemic cost of compliance and enforcement depended on the adaptability of the regulated community, a factor that the precise wording failed to anticipate.

Infrastructure planning offers another illustration. City engineers often widen arterial roads to alleviate congestion, measuring success by the increase in lane‑kilometers or the reduction in measured delay during off‑peak periods. The intervention ignores the well‑documented phenomenon of induced demand: additional road capacity encourages more vehicle trips, ultimately restoring or exceeding previous congestion levels. The engineers optimized a local geometric variable—road width—while the system‑level performance depended on travel‑demand elasticity, a feedback loop that the initial metric did not account for.

Across these examples the common structure is identical: an actor identifies a variable that is cheap to measure or manipulate locally, alters the system to improve that variable, and then observes that the expected gain fails to materialize because the variable is only a proxy for the true performance determinant. The actor’s mental model treats the component as if it were decoupled from its environment, a assumption that holds only in a narrow, artificial setting (e.g., a microbenchmark, a laboratory strain, a static cost sheet). When the component is re‑embedded in its natural context—disk I/O, material cycling, ecological variability, market dynamics, judicial interpretation, traffic flow—the coupling reasserts itself and the local improvement is nullified or even reversed.

The persistence of this pattern suggests that the underlying cognitive shortcut is robust: humans favor immediate, tangible feedback over delayed, systemic feedback. A change that yields a visible reduction in stack‑frame size or a drop in hourly wage is immediately legible; the accompanying rise in latency, failure risk, or hidden cost is often invisible until after the change has been deployed. Because the feedback loops that expose the systemic cost are longer, noisier, or mediated by other actors, they are routinely undervalued in decision‑making processes.

Recognizing this tendency does not require abandoning local optimizations altogether; it requires making the coupling explicit. In the case of Git, a realistic performance model would weigh the cost of hashing against the cost of moving the hashed object to or from storage, perhaps by instrumenting real workloads rather than relying on microbenchmarks that isolate the CPU step. In engineering, fatigue analysis must accompany pressure calculations. In biology, life‑history models must balance fecundity with survivorship curves. In economics, total‑cost‑of‑ownership models must incorporate transaction and coordination expenses. In law, impact assessments must anticipate evasion strategies. In traffic engineering, demand models must incorporate elasticity feedback.

When the coupling is made visible, the temptation to chase a locally attractive metric diminishes, and design choices begin to reflect the true constraints of the system. The Git debate, therefore, is not an isolated quirk of a version‑control tool; it is a concrete instance of a broader design fallacy that appears whenever a subsystem is tuned in isolation from its environment. Understanding the mechanism allows us to anticipate similar mistakes in other domains and to institute checks that force designers to confront the systemic costs that their local improvements ignore.

Was this worth your time? yesflatno

Sources & further reading