Designers force users to label data for training
When a comment on a developer forum complained that “having to provide names and tag every item as I go is just more work” for a GPT‑6–powered bill‑splitting bot, the grievance was not about a single buggy feature. Designers of AI‑mediated services embed mandatory data‑collection steps into the user workflow, making the extra labor demanded of the user the hidden price of the service.
The immediate manifestation of this pattern is a mismatch between the effort a user must expend and the benefit the AI returns. In the cited discussion, the user contrasted a specialized app that forces the user to tag each line‑item on a receipt with the simplicity of “$total divide by 5, I had more, let me chip in an $10.” The AI‑driven interface promised an “Intelligent UI for everyone,” yet the user’s ordinary arithmetic sufficed. The same tension appeared when the user noted that a UI explaining “parts of a bike” was unnecessary because a plain Google image search already supplied “tons of useful images instantly.” The AI service’s value proposition hinged on performing a task that could be accomplished with far less friction by existing tools, but the service’s design required the user to invest additional cognitive and mechanical effort to annotate the problem for the model.
At the structural level, two incentives intersect. First, the provider of the AI service seeks to capture user‑generated data that can improve the model, demonstrate usage, and create lock‑in. Second, the user seeks to minimize the steps required to achieve a goal. When the provider’s design forces the user to translate an informal intention into a structured representation—by naming each expense, tagging each bike component, or enumerating travel preferences—the provider extracts a richer training signal, while the user pays a hidden labor tax. The transaction fails when the marginal benefit of the AI’s output falls below the marginal cost of the user’s extra work.
This dynamic is not a novelty of the 2020s. In medieval Europe, the Worshipful Company of Goldsmiths required every piece of precious metal to bear a hallmark, a distinctive stamp recorded in the 13th‑century statutes of the City of London. The hallmark served the guild’s interest: it authenticated the work, deterred fraud, and generated revenue from the stamping process. For the craftsman, however, the requirement added a step that consumed time and resources, especially for small orders where the stamp’s value was negligible compared to the labor of creating the object. The hallmark persisted because the guild’s incentive to enforce quality and collect fees outweighed the craftsman’s aversion to the extra effort, even though many customers would have been satisfied with an unmarked piece.
A comparable friction emerged in the United States during the patent‑medicine boom of the late nineteenth century. Manufacturers such as Dr. Kilmer’s Swamp Root advertised miraculous cures, but to sell the product legally after the 1906 Pure Food and Drug Act, they were compelled to provide detailed ingredient lists and dosage instructions on the label. The regulatory requirement forced producers to invest in laboratory analysis and label design, costs that were ultimately passed to consumers. Many buyers, however, were indifferent to the precise composition; they purchased the tonic for its promised effect. The regulatory data‑collection step survived because the state’s incentive to protect public health and to generate oversight revenue outweighed the inconvenience imposed on both sellers and buyers.
In the early days of personal computing, spreadsheet software such as VisiCalc (1979) and later Lotus 1‑2‑3 (1983) required users to arrange data in a rectangular grid before any formula could be applied. The grid structure was essential for the program to reference cells by address, but it forced accountants and analysts to reformat existing ledgers, a task that could be more quickly performed with pen and paper for modest datasets. The developers’ incentive was to create a universal, machine‑readable data model that could be leveraged for complex calculations and for later data‑mining; the users’ incentive was to reduce manual transcription. The persistence of the grid format, despite its friction, illustrates how a design that maximizes the provider’s ability to process data can dominate even when it imposes extra work on the operator.
Credit‑scoring systems in the 1980s displayed a similar pattern. Early FICO models required applicants to submit exhaustive income statements, asset inventories, and employment histories to generate a numeric score. The data‑rich input improved the statistical robustness of the model and allowed lenders to differentiate risk with finer granularity. For many borrowers, especially low‑income individuals, the paperwork represented a prohibitive barrier, and the cost of gathering the documentation often exceeded the benefit of a marginally better interest rate. The lenders’ incentive to obtain comprehensive financial portraits persisted, reinforcing a system where the user‑effort tax was built into the credit‑access pipeline.
The modern AI assistant landscape replicates these historical precedents. Companies such as OpenAI, Google, and Anthropic release models that excel at interpreting structured prompts, yet they frequently expose developers to the temptation to ask users to pre‑process their intent. A “bill‑splitting” bot that asks the user to enumerate each participant, assign a label to every line‑item, and confirm the total before the model can compute shares mirrors the spreadsheet’s demand for a tidy grid. A “bike‑parts” explainer that requires the user to upload a photo, then manually tag “chain,” “crank,” and “derailleur” before the model can generate a description, reproduces the hallmark’s extra stamping step. In each case, the AI provider harvests a richer, labeled dataset that can be reused for future training, while the user shoulders the effort of converting an informal request into a formal schema.
The coupling failure becomes evident when the service is offered to a broad public that does not share the provider’s data‑centric incentives. The comment’s author observed that “AI companies are still searching for the killer idea that will attract the general public (beyond ‘be my AI bf’ or ‘better Google’).” The search for a “killer idea” reflects an attempt to find a task where the user‑effort tax is justified by a unique AI capability—such as interpreting ambiguous natural language or performing multi‑modal reasoning—that cannot be matched by a simple search or calculator. When the task does not meet that threshold, the extra user labor remains unjustified, and adoption stalls.
The phenomenon also surfaces in the design of voice assistants. Early iterations of Apple’s Siri and Amazon’s Alexa often required users to phrase commands in a rigid syntax (“Alexa, add milk to my grocery list”) rather than allowing free‑form speech (“I need milk later”). The rigidity was introduced to collect clean utterance‑intent pairs for model improvement. Users quickly learned to adapt their language, but many abandoned the feature for more direct methods (e.g., writing a note). The voice‑assistant market has since shifted toward “natural language understanding” that tolerates variability, acknowledging that the user‑effort tax of strict phrasing hampers widespread use.
Even outside software, the same mechanism appears in infrastructure. In the United Kingdom’s 19th‑century railway expansion, passengers were required to present a handwritten “ticket‑to‑board” that listed destination, class, and time, which the railway company used to compile detailed travel statistics. The ticket’s format forced travelers to spend minutes filling out a form that a simple paper stub could have replaced. The railway’s incentive to gather granular usage data for pricing and capacity planning outweighed the inconvenience imposed on passengers, and the practice persisted until automated turnstiles eliminated the need for manual entry.
Biology offers a further illustration. In the field of ecological monitoring, researchers have long asked citizen scientists to record observations using predefined taxonomic codes. The requirement to select the correct species from a drop‑down list yields high‑quality data for longitudinal studies, but it imposes a learning curve on volunteers who might otherwise note “many small birds” in a notebook. The scientific community’s incentive to amass standardized datasets has sustained the practice, despite the friction for participants.
Across these domains, the causal chain is identical: a provider defines a data‑rich interface, obliges the user to supply labeled inputs, extracts the resulting dataset for model refinement or analytics, and thereby enhances the provider’s product or market position. The user, however, evaluates the interaction on the basis of immediate utility and effort. When the extra labor does not translate into a perceptible gain, the service fails to achieve mass adoption.
The persistence of this pattern hinges on a feedback loop. Each interaction that yields a labeled example improves the model, making the service appear more capable in subsequent demonstrations. The improved capability, in turn, justifies the continued requirement for user‑provided structure, because the provider can now claim higher accuracy or broader coverage. The loop can be broken only when the provider either (a) discovers a task where the AI’s unique reasoning outweighs the labor cost, or (b) redesigns the interface to infer the needed structure from raw input, thereby removing the user‑effort tax.
Historical attempts to break the loop often involved external regulation. The 1906 Pure Food and Drug Act, by mandating ingredient disclosure, forced manufacturers to bear the cost of data collection, not the consumer. Similarly, the 1970s U.S. Federal Communications Commission’s requirement that telephone carriers provide “plain‑language” billing statements reduced the need for users to decode cryptic itemizations. In both cases, the regulator shifted the burden of data provision away from the user, thereby eliminating the user‑effort tax in that context.
In the AI domain, a comparable regulatory approach could require that models accept unstructured natural language without demanding explicit tagging, or that they disclose the data‑collection purpose of any mandatory user input. Absent such external pressure, market forces alone have proved insufficient to eliminate the tax, because the provider’s short‑term gains from richer training data outweigh the long‑term cost of reduced user adoption.
The present incident—an AI bill‑splitting bot that asks users to label each expense—thus exemplifies a universal mechanism: a design choice that makes user effort the implicit price of data acquisition, persisting across centuries whenever a producer’s incentive to harvest structured information collides with a consumer’s desire for effortless outcomes. The mechanism survives the disappearance of any particular product, any specific AI model, or any contemporary debate.
When the next generation of conversational agents is released, the same question will arise: will the interface ask the user to “provide names and tag every item as I go,” or will it infer the needed structure from the raw utterance, thereby removing the hidden tax? The answer will determine whether the pattern continues to shape the adoption curve of AI assistants or finally yields to a design that respects the user’s labor budget.