The watchdog that rewards alerts, not accuracy
A federal judge labeled Flock’s vehicle‑mounted camera network an indiscriminate mass‑surveillance system, while a pedestrian reports that repeated security‑vehicle patrols provoke him to make obscene gestures, which the patrols then log as suspicious behavior. This exchange demonstrates a mechanism in which a monitoring apparatus is incentivized to produce alerts, causing it to treat neutral actions as threats and provoking the watched to emit counter‑signals that are taken as further evidence.
The mechanism begins with the actor that receives a reward for each alert it generates. In the case of Flock, the system’s operators gain visibility, contracts, or internal metrics whenever a vehicle is flagged for unusual presence; in the street patrol, officers receive commendations, overtime eligibility, or performance scores for each incident report they file. Because the reward depends on the count of alerts and not on their accuracy, the actor has no incentive to verify whether the observed anomaly corresponds to a genuine risk. Consequently, the actor logs any deviation from the expected baseline: a car lingering at an intersection, a pedestrian pausing to tie a shoe, or a hand raised in a gesture.
The second actor, the subject of observation, learns that any noticeable behavior increases the chance of being recorded. To express dissent or simply to reclaim a sliver of autonomy, the subject adopts a conspicuous counter‑signal that is cheap to produce but highly visible to the watcher: an obscene gesture, a deliberately slow walk, or a repeated glance at the camera lens. This counter‑signal is intelligible to the subject as a symbolic act, but the watcher, lacking context, interprets it as a potential threat because its detection algorithm or training manual flags abrupt motions, prolonged stares, or non‑standard gestures as precursors to wrongdoing. The watcher then logs the gesture as another alert, reinforcing its own reward cycle.
The coupling between signal and intent is broken: the watcher assumes that the observed motion carries the same meaning it would in a hostile context, while the subject knows the motion is expressive, not preparatory. No corrective feedback arrives to tell the watcher that the logged gesture was benign; the system has no mechanism to compare the alert with independent evidence of wrongdoing. The subject, seeing that the gesture has been recorded, may repeat or intensify it, confident that the watcher will continue to reward the alert regardless of its veracity. The loop thus runs on positive feedback: more alerts generate more rewards, which encourage more permissive logging, which provokes more conspicuous responses, which generate yet more alerts.
This same logical structure appears in unrelated domains when a detector is rewarded for hits without penalty for false alarms. Email spam filters, for example, often score highly on catch‑rate while ignoring the cost of false positives. When a filter marks a legitimate newsletter as spam, the sender may alter the subject line to include all‑caps words, excessive punctuation, or deliberate misspellings—tactics that raise the filter’s spam score. The filter then flags the revised message again, registering another hit and reinforcing the rule that unusual typography equals spam. The sender’s adaptation is not an attempt to evade detection for malicious content; it is a response to a system that punishes benign variation.
Medical diagnostics that reward sensitivity exhibit a parallel pattern. A screening test that receives funding or prestige for each positive case will lower its threshold to capture more true disease, inevitably labeling healthy individuals as positive. Those individuals, informed they are at risk, may pursue follow‑up procedures that carry their own risks or seek treatments that alter biomarkers. The altered biomarkers can then be read as further evidence of disease, prompting additional testing and expanding the pool of positives. The test’s reward structure never penalizes the healthy‑person false alarm, so the cycle continues.
Financial fraud detection operates under a comparable incentive. A bank’s fraud‑prevention team earns bonuses for each transaction blocked as suspicious. To maximize bonuses, analysts may flag any transaction that deviates from a customer’s typical pattern—such as a sudden large purchase, an overseas transfer, or a series of small payments. Customers, aware that their ordinary spending triggers alerts, begin to split purchases across multiple cards, use prepaid vouchers, or conduct transactions at odd hours to avoid the pattern. These avoidance maneuvers look exactly like the behaviors the fraud model was trained to flag as illicit, producing more blocked transactions and more bonuses. The model never receives a penalty for blocking a legitimate payment, so its sensitivity drifts upward without bound.
In biology, the adaptive immune system can fall into the same trap when its clonal‑selection process rewards receptors that bind strongly to any molecular pattern. Somatic hypermutation generates B‑cell receptors with increasing affinity; those that happen to bind self‑molecules receive survival signals because the assay does not distinguish foreign from self. The self‑reactive cells proliferate, releasing autoantibodies that damage tissue. The damage releases additional self‑antigens, which further stimulate the self‑reactive clones, creating a runaway autoimmune response. The system’s reward—cellular survival upon binding—does not subtract for self‑reactivity, allowing the loop to amplify.
Historical episodes show the same dynamics when authority gains from counting transgressions. The Spanish Inquisition’s Edict of Grace (1482) offered a portion of confiscated property to anyone who denounced a heretic. Informants therefore had a material incentive to name suspects, regardless of the truth of the accusation. The accused, under torture or the promise of leniency, often confessed and named others to end their suffering, providing the Inquisition with fresh lists of names that justified further investigations. Each new name increased the informant’s reward and the Inquisition’s tally of heretics, while the veracity of the charges deteriorated.
Britain’s Sus law of 1824 empowered police to arrest anyone found “loitering with intent to commit an offence.” Officers received praise and promotion for each arrest, creating a direct reward for stops. Communities subject to heavy patrolling responded by gathering in visible groups, holding meetings, or marching along known patrol routes to assert their right to public space. These public assemblies were recorded by officers as further evidence of loitering with intent, leading to more stops and more commendations. The law was eventually repealed after the 1981 Brixton riots highlighted how the reward structure had turned a public‑order tool into a source of community estrangement.
The Salem witch trials of 1692 operated on a similar logic. Magistrates gained stature and land seizures by uncovering witches; the court paid witnesses for testimony that identified alleged witchcraft. Accused individuals, under interrogation and facing execution, often confessed and implicated others to halt the proceedings, supplying the court with new names that sustained the panic. Each confession increased the court’s count of witches and the witnesses’ compensation, while the evidentiary basis for the claims grew thinner.
Across these cases the core mechanism is identical: an evaluator receives a payoff proportional to the number of signals it classifies as positive, with no counterbalancing cost for false positives; the observed party learns that any noticeable action raises the chance of a positive classification and therefore adopts conspicuous counter‑signals to express dissent, avoid attention, or signal identity; the evaluator interprets those counter‑signals as further positives, reinforcing its own reward loop. The missing element is a corrective signal that ties the evaluator’s reward to the ground truth of the observed behavior—a verification step that would break the feedback loop by penalizing inaccurate alerts.
Because the mechanism depends only on the structure of reward and observation, it can appear in any century, any technology, and any institutional setting. The nouns change—a guild’s quality mark becomes a verification badge, a patent‑medicine advertisement becomes a sponsored result, a radar blip becomes a fraud alert—but the underlying incentive to count hits without penalizing misses remains. When that incentive is present, the system will inevitably generate its own evidence of threat, and the watched will respond with signals that confirm the alarm. The only way to interrupt the cycle is to alter the payoff so that accuracy, not volume, determines the evaluator’s gain.