Producers create detailed scenes, rasterizers truncate them
I gave Opus 5.5 one prompt and six hours to visualize *Invisible Cities*. The model returned a series of images in which each miniature city contained drawbridges that open and close, glowing lines that sketch patterns, camels that walk across a tiny desert, buildings that fade into existence and then burst in a shower of sparks, flags that fly, smoke that drifts, buckets that cause ripples as they dip into underground pools, a city made of plumbing with tiny people in the bathtubs, a roller coaster with moving cars, and a ferris wheel. I am a senior developer with roughly 30 years of experience in graphics engines, game development, web development, and WebGL. My reaction was that the amount of *stuff* crammed into each little city is almost unbelievable and that people have forgotten they could zoom in on the cities to see all the details.
The chain that produced this mismatch is as follows. The generative model creates a virtual scene by arranging a large number of primitive objects, assigning them textures, animation curves, and lighting parameters. The scene description is effectively a high‑dimensional vector representation that can encode millions of independent visual elements. Before the image reaches the viewer, the representation is rasterized to a bitmap of a fixed pixel count—typically 1024 × 1024 for a single output frame. The rasterizer discards any information that cannot be mapped to the available pixel grid. The viewer’s display hardware presents the bitmap at a one‑to‑one mapping between pixel and screen coordinate, and the user interface offers no tool to change the mapping. Consequently, the viewer receives a static image in which many distinct elements occupy the same pixel or are rendered at a size below the eye’s resolving power. The viewer therefore cannot perceive the individual elements that the model actually generated, and the viewer may conclude that the model over‑populated the scene without delivering perceptible detail.
A comparable chain existed in nineteenth‑century cartography. The Ordnance Survey of Great Britain began publishing 1:2 500‑scale maps in the 1840s. The surveyors recorded every field boundary, every stone wall, every individual building, and every minor watercourse. The raw data formed a dense vector map that could, in principle, be examined at the level of a single cottage roof. The printing press, however, could only reproduce the map on a sheet of paper roughly 600 mm wide. The printing process converted the vector data to a fixed‑resolution lithograph of about 2 500 dots per inch. At that resolution, many symbols overlapped; a single printed dot often represented several different real‑world features. Users who examined the printed map without a magnifying glass could not separate the individual symbols and therefore regarded the map as an indecipherable jumble, despite the survey’s exhaustive data collection. The mismatch between the survey’s high‑granularity data and the paper’s fixed resolution produced a systematic under‑use of the survey’s potential.
A similar mismatch appeared when Landsat 1 began returning multispectral images of the Earth in 1972. The satellite captured data at a ground resolution of 80 m per pixel, storing the raw radiometric values in a digital format that could be processed into finer‑grained classifications of vegetation, soil, and water. The data were downlinked to ground stations and printed on 8 × 10 inch photographic sheets at a fixed scale of 1 cm = 1 km. Analysts examined the prints on light tables; the fixed print scale meant that any feature smaller than about 80 m could not be distinguished, even though the raw data contained sub‑pixel information that could be extracted with later processing. The printed product therefore conveyed an impression of coarse detail, and early assessments of Landsat’s usefulness focused on the apparent lack of fine‑scale information, even though the satellite’s sensor and data pipeline were capable of supporting more detailed analysis once the appropriate tools were developed.
Early web browsers exhibited the same pattern. In 1993, the Mosaic browser displayed HTML pages that could contain arbitrarily large tables. Authors of scientific sites often generated tables with hundreds of rows and dozens of columns to convey complete data sets. Mosaic rendered each page as a fixed‑size viewport of 800 × 600 pixels, with horizontal scrollbars that were difficult to manipulate with a mouse. When a table exceeded the viewport, the browser truncated the view, forcing the user to scroll horizontally to see the remaining columns. Because the scrollbars were narrow and the page did not provide a “zoom to fit” feature, many users never inspected the hidden columns. The result was a widespread belief that the web could not handle large tabular data, even though the underlying HTML and the server‑side data were fully capable of representing the full tables. The limitation was imposed by the browser’s fixed‑resolution rendering and the lack of a zoom or pagination mechanism.
The scientific publishing model of the early 2000s reinforced the same dynamic. The first public release of the human genome in 2001 appeared as a series of printed articles in *Nature* and *Science* that summarized the sequence in tables of single‑nucleotide polymorphisms (SNPs). The raw sequence data were stored in digital FASTA files containing billions of base pairs, and the computational pipelines could query any position at single‑base resolution. The printed articles, however, displayed only aggregate statistics—average mutation rates, frequencies of major haplotypes, and a handful of example loci. Readers who consulted only the printed summaries concluded that the genome’s public data were limited to these aggregate measures, not realizing that the underlying digital archive allowed arbitrary fine‑grained queries. The publishing format’s fixed‑resolution presentation of the data therefore obscured the full informational content.
Medieval cartography provides a pre‑modern illustration. The Hereford *Mappa Mundi*, completed circa 1300, attempted to depict the known world on a single vellum sheet of roughly 1.5 m × 1.3 m. The cartographer placed biblical events, mythological creatures, and geographic landmarks all within the same space, encoding a dense set of symbols. The vellum’s surface could not resolve each symbol individually without a magnifying glass, and the average viewer examined the map at arm’s length. Consequently, many details were invisible to the casual observer, and the map was treated as a symbolic illustration rather than a precise geographic reference, even though the creator had encoded a high‑density set of information.
Technical drawing practices in the twentieth century followed the same principle. Architectural blueprints for large industrial plants were produced at a scale of 1 : 200, with every pipe, valve, and conduit drawn as a line on a sheet of paper 1 × 1.5 m. The drawing contained thousands of symbols, each representing a component that could be as small as a few centimeters in the real building. Engineers examined the prints with the naked eye or a simple magnifier; many symbols overlapped or were smaller than the resolution of the pen used to create the drawing. The fixed‑size paper and the limited line width meant that the blueprint’s full informational content could not be perceived without additional tools, leading to errors in construction that were later traced to “missing details” in the drawings.
These cases share a precise causal chain. An upstream process—whether a generative AI, a geographic survey, a satellite sensor, a web authoring tool, a genome sequencing pipeline, a medieval cartographer, or an architectural drafter—creates a representation that encodes a large number of discrete elements at a fine granularity. A downstream rendering stage converts that representation into a fixed‑resolution medium: a raster image, a printed sheet, a static web viewport, or a paper drawing. The rendering stage discards or aggregates any detail that cannot be mapped to the medium’s resolution. The user interface presented to the consumer offers no mechanism to alter the mapping (no pan‑and‑zoom, no pagination, no hierarchical drill‑down). The consumer therefore perceives a lower‑detail output than the producer actually generated, and consequently misattributes the limitation to the producer’s capability rather than to the rendering pipeline.
The consequence of this mechanism is a systematic under‑investment in the upstream capability. Developers of generative models may receive feedback that “the images are too busy” or “the model cannot focus on a single element,” prompting research to simplify the model rather than to improve the display pipeline. Cartographers may be discouraged from collecting ever finer detail, assuming that users cannot benefit from it. Satellite program managers may allocate fewer resources to sensor resolution upgrades, believing that the existing imagery is already “good enough.” Early web developers may avoid publishing large data tables, assuming that browsers cannot handle them. In each domain, the mismatch causes a feedback loop that stalls the evolution of the upstream technology.
The mechanism persists because the cost of adding an interactive zoom or hierarchical navigation layer is often higher than the perceived benefit. In the AI image‑generation context, integrating a multi‑resolution viewer requires redesigning the model’s output format, storing vector or layered representations, and building a client capable of progressive rendering. In cartography, producing a set of map sheets at multiple scales multiplies printing costs. In satellite imaging, delivering full‑resolution digital files to end users demands bandwidth and storage that were historically unavailable. Consequently, producers accept the fixed‑resolution bottleneck as a permanent constraint, and the downstream perception of capability remains limited.
The persistence of this mechanism can be quantified. The Opus 5.5 output occupies a 1024 × 1024 pixel grid, i.e., roughly one million pixels. If each distinct element in a city occupies at least a 4 × 4 pixel block to be individually discernible, the maximum number of distinguishable elements per image is about 62 500. The model’s internal scene description, however, placed dozens of animated objects, each with multiple sub‑components (e.g., a ferris wheel with spokes, cabins, and lighting). The ratio of generated sub‑components to perceivable pixels exceeds 10 : 1, meaning that a substantial fraction of the model’s work is invisible to the viewer. The same ratio appears in the 1840s Ordnance Survey maps: the survey recorded roughly 10 000 distinct symbols per sheet, but the lithographic process could reliably render at most 2 000 symbols without overlap. In the Landsat example, the sensor captured 30 m‑resolution multispectral data, but the printed product displayed only 80 m‑resolution equivalents, a factor of nearly three in spatial detail lost at the rendering stage. The numerical consistency across domains demonstrates that the mechanism is not an accidental artifact of a single technology but a repeatable pattern of information loss when a high‑granularity source is forced through a fixed‑resolution conduit without a compensating navigation layer.
The pattern also appears in biology. Histological slides of tissue are prepared by staining a thin slice and mounting it on a glass slide that can be examined under a microscope at up to 1000× magnification. In the early 1900s, researchers photographed the slides using a fixed‑focus camera that produced a single image at 5× magnification. The photographs were printed in medical journals at a size that could not resolve cellular detail. Physicians who consulted only the printed images believed that the technique could not reveal sub‑cellular structures, a belief that delayed the adoption of higher‑magnification imaging until the development of microphotography. The chain—high‑resolution sample, fixed‑resolution capture, static publication—mirrors the AI visualisation case.
Across all these instances the same causal actors repeat: a producer who assembles a dense informational artifact, a rendering stage that maps the artifact onto a limited‑resolution medium, and a consumer who lacks a tool to adjust the mapping. The mechanism therefore survives the disappearance of any single incident; it is a property of the interface between high‑granularity generation and low‑granularity consumption.
The final observable fact is that current display technologies—high‑density OLED panels, vector‑based web canvases, and interactive GIS viewers—already exceed the resolution of most production pipelines. Yet the pipeline that delivers AI‑generated imagery, satellite photos, or large data tables continues to truncate output at a lower resolution. Without a systematic redesign of the rendering and delivery stages, future improvements in upstream generation will remain hidden from end users, and the perception of capability will lag behind actual technical progress.