The Friedrichstrasse Deepfake

Some thoughts on representation: forced perspectives, sampled beats, marked sentences and the long habit of blaming the machine
In summer 2026, the European Union began mandating labels on realistic images that are produced with artificial intelligence. Legal advisors for architectural publications think that digital renderings fall under this rule. The same legislation requires generated text to contain a specific mark in the writing, and it states that providers must contractually prohibit users from removing it. The two issues, which appear to be on different levels, are deeply interconnected.
What follows is an argument about what an artificial representation has ever claimed to be. The answer runs back seven centuries, and the shortest way into it is a photograph that Mies van der Rohe altered by hand in 1921.
1. The Charge
In late 1921, Berlin’s Turmhaus-Aktiengesellschaft hosted a design competition for Berlin’s first skyscraper to be built in a prominent triangular plot bounded by the Spree River, the Friedrichstrasse shopping street, and the Friedrichstrasse train station. The brief allowed six weeks and asked for simple hand-drawn perspectives from two designated viewpoints. Among the more than 140 entries, the best-known project was the one submitted by Ludwig Mies van der Rohe, under the motto “Wabe”, honeycomb: Mies moved both viewpoints far back, commissioned two large photographs of the street from the new positions, and drew the building in charcoal directly onto the prints. He was the only competitor who did any of it, and a serious jury would likely have disqualified him right away for failing to comply with the contest guidelines. For the record, he didn’t win the competition.
The result has been called a photomontage for most of a century. It is not. Nothing is cut and nothing is pasted. The tower is drawn into the silver halide, a hand-controlled veil laid over a document, and the seam between the invented and the recorded never occurs. Pablo Gallego-Picard, writing in the “Rendering” issue of CLOG in 2012, calls the results drawn photographs. The operation has a name now, and a toolchain: Photoshop shipped the clone stamp in 1990, the healing brush in 2002 and content-aware fill in 2010, each of them keeping a photograph as the ground and replacing part of it with something that was never there. Generative fill does the same job with a diffusion model behind it and a sentence as the instruction. What has improved since 1990 is the quality of the seam. Mies closed his by hand in 1921.


Gallego-Picard demonstrates a more serious thing as well. Mies’s elevation and his ground floor plan do not agree with each other, and neither agrees with the images. Model the volume from the drawings, put a camera at the two viewpoints Mies chose, and a different tower appears: correct in its descriptive geometry, and slack (images 3, 4). The compelling building we know from having seen it featured countless times in architectural magazines and history books as a benchmark of the discipline, exists in the two original charcoal views and nowhere else. The project that could have been built is another building, and a way duller one.
Pablo Gallego Picard and I were unaware of each other’s existence until that issue of Clog, in which I was also published with a short essay of my own, “Real VS Plausible”, also written one decade before public diffusion models existed. My argument then was that we manipulate renderings today as we forced perspectives yesterday and axonometrics before that. Without having discussed it beforehand, Gallego-Picard and I were both expressing the same idea, which is obvious but not emphasized enough.
The technique was six centuries old when Mies used it.
2. The Augmented View

Giotto finishes the Scrovegni Chapel in 1305. If you look closely at the sides of the chancel arch that frames the apse from the main nave, he paints two small rooms that are not real: in both, ribbed vaults, a window with a view into nothing, a hanging lamp. Simulated architecture inserted into real architecture, in one of the earliest coherent illusions of depth, a century before the rule was written down. Nothing on that wall exists, yet that extra space seems believable and the eye accepts it.
The formal rules were established a century later. Brunelleschi carried out experiments around 1420, and Alberti recorded them in 1435. Afterwards, a flat surface could show spaces and volumes that did not exist in the real world. The eye perceives them as solid and logical when the viewer stands in the proper position. After that, the flat surface was permanently altered. Entire worlds that had never been built became available, credible, and perception-wise consistent.
Two Roman buildings employed this technique in different ways.

In 1652, Borromini began to skillfully manipulate physical space in this way, shaping how it is perceived. He designed and built a colonnade at Palazzo Spada together with the mathematician Giovanni Maria da Bitonto, who supplied the required calculations. The vanishing point, normally located at infinity, is set at a fixed spot a short distance from the viewer. All elements are directed toward this point. The floor rises sixty centimeters, the barrel vault drops, the walls become narrower, and the columns reduce in height and diameter as they recede. This design speeds up the perspective, making a nine-meter passage appear as a monumental gallery exceeding thirty meters. No part of the effect relies on painting. The physical space is distorted until the eye perceives a different building. That is the purpose of trompe l’oeil. It is a manipulated representation that is believable enough to be trusted. It provides a sense of scale that the physical site could not accommodate.


In 1685, the Jesuits at Sant’Ignazio in Roma could not construct their dome due to a dispute with donors. They hired Andrea Pozzo to paint a fake dome on a circular flat canvas. A marble disc in the nave floor indicates where the viewer must stand for the illusion to work: if one moves away from that spot, the dome appears distorted. It is an image of a building part that was never built, placed where that building part was intended to be seen.
Both buildings function by selecting a specific viewing point in advance, as designed, and both fail to work as intended when the observer moves. Since manipulation was the goal here, removing it vanishes the interest to those spaces.
The architecture profession answered this with a defensive conclusion. The extravagance of the Baroque era had preferred a single viewpoint. That was supplanted by a demand for objectivity that arose together with rationalism and the sciences. In the next century, architects were expected to employ orthogonal projection, with precise angles and uniform lines that can be checked against a physical structure. Perspective was seen as a personal viewpoint and was assigned to painters and stage designers.
Later, the exact science of projection was further developed and crafted for military and ballistic use.
3. The Weapon
In 1765 at Mézières, the French draftsman Gaspard Monge was tasked with positioning guns in a fortress, to shield defenders from external fire. The usual arithmetic method was slow, so Monge applied geometry to obtain a quick solution. His commandant first rejected the result due to its speed, but after further inspection the method was classified: descriptive geometry stayed a French military secret for roughly thirty years because precise spatial control gave a clear ballistic advantage. Monge was not allowed to publish his work until 1799.

A single discipline deceived an eye in Rome and also affected the use of artillery in France. The same reasoning is applied in modern — or rather, these days, classic — rendering engines. These programs compute how light strikes surfaces and project the outcomes onto a flat plane employing the methods Gaspard Monge established. So even governments have historically tracked who holds this geometric knowledge. Its honesty was never in question, because every step of it is science that can be demonstrated.
The line stops there. Diffusion and generative models work in a different way: they hold no 3D scene, no camera and no viewing position, and they run no calculation based on any of them. They arrange pixels to match the statistical patterns of photographs taken by people who had all three. The result is estimated rather than solved, which is why long straight lines and reflections are where these models struggle. Francesco Borromini set a fixed position for the viewer and designed outward from that point. A diffusion model produces an image that leads the viewer to assume a viewpoint existed, even though no such perspective was ever defined. In this case too, it is a matter of manipulated representation that is believable enough to be trusted.
Which brings us to 2 August 2026, when Article 50 of the European Union’s Artificial Intelligence Act came into force. The regulation obliges publishers of realistic images, audio, or video to indicate whether the material was generated or modified by artificial intelligence. From now on, a trivial rendering of realistic scenery with people meets the Act’s definition of a deepfake, if generated through AI. Such images must now carry a label.
Mies’s drawing corresponds to every element of that description. He used charcoal. A hand-drawn deepfake, created in 1921.
4. What the Image Claims
The EU Act asks one question of an image: was it made by a machine? It does not ask what the image claims.
Two examples from this 2026 summer show the distinction between the two issues. In late July, Google introduced a feature in Google Earth that let users overlay generated scenery onto satellite, aerial, and 3D imagery. Users were then able to download the resulting image. Within a day, researchers, jokers and also joking researchers employed the tool to produce images of a nuclear plant in Iran, a refugee camp on the Mexican border, and a bombed hospital in Gaza. Google took down the feature within twenty-four hours.

Google Earth serves often as a forensic-grade tool for verifying events. Open source investigators employ satellite imagery to verify claim accuracy because this data has historically been costly and hard to counterfeit. The possible harm is considerable for two reasons. False images can spread rapidly, and governments acquire a reason to call true ones fake, or vice versa.
An architectural rendering has never made that claim. It shows a building that does not exist, on a day that has not occurred, under weather selected for the purpose, with figures who are there to give scale. In the same 2012 essay on Clog I described the profession’s own understanding of this: we allow others to believe that rendering is the most objective way to represent architecture, and we all know it is not. The image asks to be found plausible. It does not ask to be believed.
Article 50 puts both images in the same category. A satellite photograph provides evidence of existing conditions, a rendering acts as a proposal for something that does not exist. The Act regards them as the same object when a diffusion model generated the rendering. So, a view produced through “traditional” V-Ray engines can perfectly resemble a photograph yet should stay exempt from the rule, but the same view bears a label if it was “AI-generated”. Two images that make identical claims are classified according to the software used to create them. So the rule lands on the image that never asked to be believed, and adds nothing to the one that did: the company that opened the satellite case shut it in a day.
Things are even more complicated here, since the distinction “AI-generated”-or-not also vanishes during the production process. Modern rendering software employs neural networks for denoising and upscaling by default. Nearly every professional image made in 2026 includes AI elements. People working in these studios cannot pinpoint the specific stage in the workflow where the transition to AI occurs.
5. Every Instrument Was Feared, Every Instrument Was Bent
Every technology used for representation claims to offer a more accurate likeness. Within one generation, these tools are used instead to create a preferred version of reality. Printing made it possible to reproduce drawings, and architects first employed this capability to alter historical records. In the 1570 publication Quattro Libri, Andrea Palladio displayed several of his own villas with straightened lines and added wings that were never actually built. These printed plates, rather than the physical buildings, shaped architecture in England and Virginia. Photography was originally presented as a form of objective evidence, yet users were staging and retouching photographs within twenty years of its invention, a century before Photoshop.

Generative models represent the most recent development in this sequence. They are the first models capable of operating on both visual and written content. Since they can produce images and sentences, discussions of realism now apply to representation and writing at the same time. The main change concerns the level of accessibility: previously, creating a believable distortion of reality required years of specialized training. That requirement restricted the practice to individuals who had professional reasons to consider the impact of their work. That technical barrier no longer exists, and the technology of generative models improves week by week.
But also in the past, new technology has provoked a reaction. In January 1982, Barry Manilow toured the United Kingdom with synthesizers rather than a live string section. This change led to orchestral musicians losing their jobs. On 23 May, which was Bob Moog’s forty-eighth birthday, the Central London branch of the Musicians’ Union decided to ban synthesizers and drum machines. They also sought to ban any electronic device that could mimic the sound of a traditional instrument. The branch expressed concern about the West End. They feared that technicians would eventually replace musicians in orchestra pits. The NME described the union members as crazy, and synthesizer players answered by forming a competing organization. Eventually, the national union altered its stance to a request for job security measures. No equipment was ever banned.
The motion has aged into comedy. The union feared a machine that could imitate a violin. These machines were poor at imitation. Their actual value came from sounds an orchestra could not make, and from musicians who did not try to hide the electronic nature of the instrument. The Musicians’ Union attempted to ban simulation but failed to recognize the new medium.
Nine years later the second pattern appeared. On 17 December 1991 Judge Kevin Thomas Duffy opened his ruling in Grand Upright Music v. Warner Bros. with four words: thou shalt not steal. Biz Markie produced a track that incorporated ten seconds of music from Gilbert O’Sullivan. Warner issued the album prior to approval of the clearance request. Judge Kevin Thomas Duffy placed an injunction on the record and suggested that the case be pursued criminally. Prior to this ruling, many producers believed that short musical fragments could be used without legal permission. That belief turned out to be incorrect. A study of Billboard hip hop songs covering 1988 to 1993 demonstrates the impact of this decision. The average count of samples per song fell, and producers altered the way they employed the remaining samples. De La Soul paid 1.7 million dollars to resolve a lawsuit with the Turtles.

Sampling continued, yet the practice became restricted to those who could afford it. Clearance fees are fixed expenses that remain unchanged regardless of a business’s size. These expenses proved most challenging for independent producers lacking the backing of a major record label. Other artists started employing interpolation. They hired session musicians to re-record particular musical fragments. This let them license the underlying composition while sidestepping the fees tied to using the original master recording.
The commandment thou shalt not steal has had a second life: it’s the charge now laid against generative models, trained on professional work that was never cited and never paid for. This objection is the same one the Turtles raised versus De La Soul, although it now occurs on a larger scale.
So there are two ways for this to go wrong, and both are already visible. A ban that focuses on simulation ignores the actual purpose of the technology. Additionally, requiring a fee for compliance lets the practice continue while restricting it only to those who can afford to pay.
6. The Verified View
British planning systems already differentiate image types in a manner that the Act does not.
An Accurate Visual Representation — AVR — is a photomontage derived from a survey: the camera position is captured with GPS, and twelve to fifteen points within the photograph are surveyed, so the model is subsequently aligned to those data. A method statement accompanies the image. The London View Management Framework classifies them from AVR0, displaying only location and size, to AVR3, displaying full materials. Planning authorities select the grade according to the requirements of their decision. Studios also generate unverified CGIs, labeled accordingly and meant to illustrate the general intent of a project.
The idea of sorting images by their reader is older than the framework. Filarete’s 1460s treatise identifies three types: a sketch for the designer’s own testing, a looser drawing for the client that may employ perspective, and a scaled drawing for the builder. The persuasive drawing is labeled as such and provided to the person who needs to be convinced.
The distinction is older than the software in the other sense too. Masaccio positioned the viewpoint of his Trinity at eye level for a viewer in Santa Maria Novella. The painted chapel seems to open into the wall because the geometry mirrors the viewer’s actual position. Pozzo’s dome employs the same technique for the opposite effect and indicates the precise spot on the floor where the illusion functions. These represent two different image statuses based on their purpose.
The AVR system does not rely on the method of pixel generation. A verified view stays valid regardless of generative models because the survey data either confirms the image or does not. Twenty years of development produced four specific grades, whereas the Act relies on a single binary choice.
7. The Undetectable Difference
So, there is a distinction the Act cannot draw. An image may be almost automatically generated and then shipped. Alternatively, a model can sit in the orchestra while a human director makes all the important decisions. These are distinct acts, and no label separates them.
Photorealistic render engines appeared in the mid-nineties and required a skilled human to operate them. No one requested those images to be labelled. The discipline functions by threshold, argued case by case, and is enforced by clients and editors who have seen enough to recognize when a studio has gone too far.
Laziness leaves no trace within the file. A detector discovers the signature of a generator, which indicates where the pixels originated. It cannot count, or qualify, decisions. An image edited over three days still bears the fingerprints of the model that opened it. Slop pushed through sufficient compositing loses them. The machine cannot see the relevant difference, and the Act is asking the machine.
What remains is the author’s word, which is also what remained in 1921. No inspection of the Friedrichstrasse images reveals how much of them Mies decided, and how much followed from the tool, from charcoal worked over a photographic print and the limits that imposes. We know this because he signed them. The Act swaps that test for a question of whether a machine was involved. But now a machine participates in every stage of every workflow in many industries.
8. Inside the Instrument
The European Commission’s Code of Practice on Transparency of AI-Generated Content took effect on August 2, 2026, the same day of Article 50. Its requirements are narrowly defined. Any generated text exceeding about 150 words must bear a mark, providers must contractually prohibit users from removing that mark, and the mark must persist through screenshots, OCR, and translation. The Code was written with several distinctions: between generation and assistance, between different markets, and between various uses.
The technique has a broad cost. Prose lacks a metadata layer in which a label could be concealed. Words themselves constitute the file. Applying a mark to text requires selecting different words. At each decision point, candidates are divided by a secret key into two groups, and the model is steered toward one; after enough words, the bias can be detected by the holder of the key. Google DeepMind demonstrates the method using a sentence about tropical fruits that resolves to bananas, and claims the reader will not notice. The method renders it permanently unanswerable whether pineapple would have been the better word.
The paradox resides in Anthropic’s implementation. It is a blanket, model‐level approach that exceeds the Code’s requirements: it covers every Claude model in every market, even in jurisdictions with no legislation, justified by the claim that regional scoping is not yet available. It also applies to text that the model only edits, meaning a hand‐written paragraph submitted for proofreading could be marked as machine‐generated. It extends to private conversations that no third party will read.
A flawed rule already causes one kind of harm. When a flawed rule is met with over‐compliance, it causes a way different, larger harm. The Code requested a simple mark on published output. The implementation removes those distinctions, expands its scope to countries that requested nothing, and shifts the cost onto the instrument itself, landing on the words that provided the most accurate answer to the question.
The analogy compares it to a calculator that must produce results that are occasionally, imperceptibly wrong, off by the last digit, on a schedule known to the manufacturer, allowing a long enough series to be traced back to the device. No profession would have tolerated such a condition. Arithmetic includes a bell that rings when an answer is incorrect. Prose lacks a bell, a fact concerning detection rather than damage. The providers concede this without appearing to notice: Anthropic says the mark is omitted when a single correct token exists, and is applied more lightly to code because code must be exact. The exception gives the argument away. The mark costs something, and the provider has chosen where that cost is allowed to fall: on prose, where nobody can prove the word was wrong.
Each distortion in this essay has served someone: Pozzo’s fake dome served a church that could not afford one, and Mies’s charcoal served a project he wanted to win; both are part of the work and can be judged as such. The watermark serves a compliance report, is applied to every sentence regardless of whether it needed it, and the author responsible for the result never chose it.
James Padolsey identified the offended principle: treating assistance as suspect only once the tool can compose a whole sentence imposes, in his words, a moral premium on difficulty itself. John Gruber is more concise, calling the EU Code as it is: “red‐tape nanny‐state pipe‐dream nonsense”.
9. What Actually Rots
In July 2024 Shumailov and colleagues reported in Nature that training a generative model recursively on the output of generative models causes irreversible damage, with the distribution tails disappearing first and the centre narrowing with each iteration. The phenomenon appears in language models, variational autoencoders and Gaussian mixtures, indicating that it is a property of the procedure rather than of any specific architecture.
Volume decides the rest. Slop can be generated in seconds at no cost, unsupervised and at any scale, whereas material suitable for training is produced slowly, costs a lot, and is created by people piece by piece. A corpus does not evaluate a sentence by its value, it merely counts, so sheer abundance prevails and the balance shifts further with each generation. That is Darwinian in the literal sense: the entities that survive are those that replicate most easily, and the mechanism does not demand any quality.
Idiocracy portrayed that scenario in a fictional future; the situation we face is quieter, a gradual decline of the median until even simple queries return degraded results.
The web serves simultaneously as the training corpus and as a dumping ground, and every published item feeds into the next model. That is the sole area where disclosure could provide measurable benefit, and it is not the purpose for which either tool was created. EU regulations claim a control on what the public sees, while the crawlers operate unsupervised on sources whose generated material has already become contaminated beyond philological accounting. The binary label provides no usable information to a reader: it tells how the image was created at the moment when the relevant question is what the image asserts, and it leaves all upstream aspects unchecked.
I have never objected to the AI itself when it comes to slop. Simon Willison coined the term in 2024, modeling it after spam: material created without review and sent to someone who never requested it, with not all generated material meeting the criteria, just as not all promotional mail counts as spam. The definition depends on review and ownership rather than on mechanism, and Willison’s own rule is that he attaches his name to whatever he publishes, meaning no detector can ever resolve it. Someone publishes, another person asks what was reviewed and who is responsible, and the response is debated publicly. Last Christmas an artist credited on an Apple promotional image was asked if he had used generative AI and declined to answer; critics spent a week discussing the non‐denial, and the work was harmed by the reticence rather than by any label. The process is slow and cannot be automated, and it yields criticism: a collective sense of what a field will accept, adjusted case by case by practitioners and open to challenge by anyone willing to argue.
Curating a training corpus also counts as criticism, and in this case the argument works against me. A model that degrades on recursive data must be fed selectively, meaning someone chooses what goes in. This is already how it functions. Fewer than ten companies in the world train models of this size, among them OpenAI, Google, Anthropic and Meta, and each makes that choice upfront, at scale, owing nobody an explanation. None of them publishes what went into the corpus, so nobody outside the company can contest the selection, and it stays proprietary.
The cure then produces the same effect as the disease: contamination removes the tails of the distribution, as does taste, so a model fed only on what is already agreed to be good returns a prestigious average and nothing else. Clement Greenberg drew the line between avant‐garde and kitsch in 1939, in a magazine, under his own name, and spent the rest of his career defending it. The AI watermark repeats the pattern, in negative: both keys sit with the provider, so nobody outside can confirm or deny where a text came from. Someone will always be judging this material. Whether the rest of us can argue with the judgement depends on the format in which it is recorded.
10. Coda
So the EU passed a law to harness Artificial Intelligence. It made disclosure compulsory for realistic images produced with an AI system, and asked that generated text carry a mark readers cannot remove, on the reasoning that a public told what a machine had touched would judge it better. The judgment it aims to protect is not the one in question, the mark can be read only by the company that created it, and the corpus from which the entire issue originates remains untouched.
Brussels has employed similar measures previously: the Digital Markets Act demonstrated that when compliance costs exceed the value of the European market, the feature is not released here. Watermarking works in the opposite direction, a rule intended for 450 million people is applied to eight billion because separating the groups is inconvenient. The more immediate risk for creative professions is less noticeable than either of those. European designers will probably flatten their images out of caution, releasing in AVR1-grade what would have been released at AVR3, while persuasive images move to jurisdictions that do not require a label: so the Act does not increase honesty but could considerably reduce ambition.
That same criticism also applies to words. A mark embedded in word choices informs a reader that a model was involved and then stops, providing no information about which passages were originally drafted, which were revised, or whether anyone reviewed the result before publication. One bit does not constitute a fact about a work, and on either side of the page that single bit is all the label conveys.
The version that is worthwhile to build is not designed for us. I think that if recursive training on generated material irreversibly degrades models, the training systems must know what they are receiving, and a machine‐readable problem requires a machine‐readable solution: a provenance record attached to the output that enumerates the operations applied in sequence and the tool used, which approximates what the C2PA content credentials aim to provide. If one generates and publishes, it records one step; if one generates, rebuilds and corrects across forty passes, it records forty steps. It cannot quantify how much of that was a human decision and should not claim to do so. Recording what was done already exceeds what we currently possess, and anyone can verify it, something the watermark does not permit: the keys are now held by Anthropic and Google, no single detector applies to every model, and services that remove the mark through paraphrase appeared online within weeks. Interpolation already done in music sampling recurs, thirty‐five years later.
If graded, it would operate like in the UK planning, assigning images from “AVR0” to “AVR3” based on how much a viewer can rely on them. Readers might disregard the scale just as they disregard an exposure value in an image file; the crawler and the archive would not be able to. A training corpus capable of distinguishing its own output from ours, and a record for anyone attempting in thirty years to determine what a document was, are more valuable than a mere warning, and a mere warning is what Article 50 created.
We have previously created sophisticated tools and employed all of them to generate culture and to forge it, often using the same hands. No one considered labeling the output of a calculator. We required the calculator to be correct and the operator to display the workings, and that arrangement has persisted for decades. It relies on a professional human signature, the sole element of this process that has ever been accountable, and the part these rules leave unchanged.
Fourteen years ago I concluded the CLOG article on this topic by quoting Schopenhauer, stating that the world is our representation, and noting that architects have always known this. It provided a comfortable stopping point, but I will not stop there again. We have been operating our own “artificial intelligence” for centuries, together with our very human algorithms, on all intellectual output, both images and text, and much of that output would meet the Article 50 flag: we imposed perspectives, straightened villas in print, sanitized competition views, and then deemed the gap between the image and reality an acceptable tolerance. A rule requiring every persuasive image to state that it is persuasive would encompass all that this profession has created, including my own work, and we would have deserved it. I would endorse that rule.
Article 50 only inquires whether a machine was present. It never queries whether the image is presented as evidence of something that exists or as a proposal for something that does not, therefore a rendering not intended to be believed receives a label whereas a verified view appears beside it without a label. Examination focuses on the production means rather than the claim, and the claim is the sole element of an architect’s image that has ever been subject to verification.
It exonerates us: seven centuries of images that promised more than the real thing, and Brussels targeted the software. Chapeau.
A rule designed to suit a process of this magnitude will eventually appear, but it will be too late, with the models already irreparably contaminated by source material that no one can qualify. And no flag able to record how short‐sighted the Union was at a pivotal moment in the history of information.
Luca Silenzi, August 2026
The Friedrichstrasse Deepfake was originally published in Bootcamp on Medium, where people are continuing the conversation by highlighting and responding to this story.