Using generative methods for artistic experiments

Hito Steyerl and Francis Hunger
Cite as
Steyerl, Hito; Hunger, Francis: "Using generative methods for artistic experiments". carrier-bag.net, 1. October 2026. https://carrier-bag.net/using-generative-methods-for-artistic-experiments/.
Import as

Introduction

Five practice reports illustrate the scope of research and studying within the Emergent Digital Media class (Prof. Dr. Hito Steyerl, Dr. Francis Hunger) at the Academy of Visual Arts, Munich. They foreground examples of generative practices using machine learning techniques, colloquially known as ‘AI’, leaving aside other diverse artistic strategies (conceptual, performative, video essay, etc.) negotiated in class. https://www.generativemedia.net

The scope of the reports by Vasilii Vikhlaev, Chloe McFadden, Guillaume Menguy, Nikita Sazonov and Otto Ostermann reaches from direct inquiry into techniques and tools to the metaphorical reflection on societies ‘AI’ phantasms. A commonality of the diverse approaches is that they probe how systems work, instead of just using the outputs as finished results. By that the artists develop a critique of the dominant narratives about ‘AI’, and they uncover the material dependencies and massive infrastructures (data, energy consumption, data centers) behind the stochastic generation of images, sound and text.

Vasilii Vikhliaev: Machine Learning and Sound as experimental field

This report addresses two registers of my work with machine learning in sound: the inductive one, meaning pre-trained models, and the deductive one, meaning physical modeling. The question behind both is how to synthesize sounds that are not present in any database. I’ll introduce three projects and deliver an outlook on future research.

The first project is Aggressor’s Tongues (2025/26), a 10-channel audio experiment of 41 min that was exhibited at the Kunstbau of Lenbachhaus, Munich. The project began with audio material of scraped recordings of Russian propaganda speech from the ongoing war against Ukraine: voices saturated with hate and domination. The machine listens, so I don’t have to. The Montreal Forced Aligner (MFA) algorithm takes an audio file and its transcript and returns time stamps for every word and phoneme, normally as a preparation step for phonetics research or for building speech datasets. I used its output as the material itself: the piece works on the phoneme grid, not on the meaning of the speech. The grid also set the segmentation of the training data for variational autoencoders (VAE), so the model could only learn units below the level of meaning. The VAEs were activated using the RAVE framework, originally developed at the Institut de recherche et coordination acoustique/musique (IRCAM).
During the process, I discarded an earlier approach with emotion recognition (openSMILE, Praat), because the results were too illustrative and literal. Instead, I focused on decoder artifacts. I displayed them rather than smoothed them out, to keep the apparatus audible. For the same reason the drifts through the latent space were deterministic, several phasors in irrational ratios, so that the trajectory never repeats. I wanted to hear and make audible the system, not the sounds it was meant to produce.

The second project is titled, vzkazy domů (2025) and was developed in cooperation with Andrea Veselá. It’s format is 7.1 surround sound with a length of 10 hours, and was developed for the former Radio Free Europe building in the frame of Public Art Munich 2025. Starting point was a short found-footage recording of a historical radio signal being intentionally jammed to interfere with its broadcast, going from signal to noise. We slowed this recording down and Andrea contributed a reactive voice interpretation. Then she listened to her own recorded voice through an earpiece and sang to it, close enough in pitch that the two voices interfere and produce interfering oscillations, called ‘beatings’ (Schwebungen). That material became the training data for the VAE. The preparation of the input was prioritized over any selection at the output, emphasizing the data: the beatings were already in the training material, as a measurable modulation, and not added posterior as an effect.

For the third project, Composite (2026, in progress), I turned a synthetization technique that involves no machine learning at all, called ‘physical modeling’. The work is based on the composite plastic coins of Transnistria, the first plastic coins in general circulation. With Modalys, a generative tool developed at IRCAM, sound is computed from the physical properties of a body (size, density, Young’s modulus). Unlike classical sound synthesis (additive, subtractive, FM) or sample-manipulation approaches, physical modeling simulates the sound-producing behavior of an instrument itself. This makes it possible to capture sonic qualities that traditional synthesis techniques reproduce only with difficulty. As part of the exploratory artistic process, I went from small data to no data. Physical modeling uses established physical laws from material sciences and acoustics, so it can compute objects that never existed. Compared with machine learning approaches, physical modeling is able to produce sound without examples. Machine learning for audio is almost always inductive, needing data, so the availability of data decides what can be produced at all. For rare, local or non-existent material that is a hard limit. I did not take this step to abandon machine learning techniques, but to run both methods side by side and hear the difference.
For the machine learning process, I am building a dataset of the real coin sounds to train a VAE on them and drive that model with the computed sounds using physical modeling: the input is encoded into the latent space and decoded again, so the output carries the statistics of the training data. The input cannot lie inside that distribution: Modalys computes an idealized body, without a surface, without the other coins, without a room. Some of the material constants I set, do not belong to any real object either. The statistical signature of a real object is imposed on an object that never existed. Input and output will be presented as pairs, not blended, to keep the model readable. The inverse is planned too: physically modeled sounds as training data, real coin sounds as input driving the model.

Chloe McFadden: Prompt Shifting

This report reflects upon my practice-based technique, prompt-shifting. It is both a practical technique for methodically probing the biases and patterns of text-to-image models, and a conceptual intervention that opposes magical technoscientistic framings.

Prompt-shifting enables artists to probe commercial text-to-image platforms without stabilizing their rhetorical claims and purported abilities of dream-making and superior accuracy. Rather than visualizing our imaginations, this technique shifts attention away from outputs towards the processes and contexts of their production. Using the same seed, the artist slowly and methodically introduces and varies the influence of certain subjects in the prompt. By noticing what shifts and emerges overtime, the artist can speculate and sense how learnt patterns are activated within generative AI models in ways that manifest visually.

This technique emerged from my broader research project that attempts to engage and disrupt the ‘magical technoscientism’ of generative AI: a fusion of magic and technoscientism that affords models the ability to simultaneously create miraculous and scientifically authoritative outputs. In commercial text-to-image rhetoric, prompts are imagined as granting users the gift of dream-making while appeals to ‘prompt-adherence’ and ‘accuracy’ position prompting as a science. Such a dissonant fusion positions models as both passive conduits of human imagination and objective authorities of visual representations. Prompting guides and practices frequently frame the misalignment of expectation and output, not as a limit of the system but as the user’s failure to conform to its representational logic. The problem is how the user asked, not what the system can do.

This redistribution of misalignment creates a magical technoscientistic perception of text-to-image models in which prompts are both magical phrases – ‘open sesame’ – and formulas: dog + beach + realistic + 35mm film – people = output image. Via this magical perception, positive and negative prompts are imagined as adding or subtracting certain influences and features from an output image. Such explanations reduce image generation to a representational exchange of words for images, obscuring the operations through which outputs arise. Thinking operationally, what does it mean to ‘add’ or ‘subtract’ an influence from the creation of an image? Prompt-shifting attends to such a line of inquiry, redirecting attention away from representational desire and towards operational curiosity. For example, methodically introducing negative and positive terms to a prompt foregrounds the operativity of the model and how terms are embedded and mutually determine the trajectory of the denoising process.

Prompt-shifting thus enables a sensing of image features not as discrete and universal parameters, but as relationally, socially and operationally situated. This situated sensing may also disrupt framings of bias as an issue that can be resolved through the acquisition of more data. Prompt-shifting both requires a reflection upon the situated conditions of production enacted by text-to-image applications (across models, society and time) and demonstrates how bias manifests visually within generative images in strange and unexpected ways.

Guillaume Menguy: Friction, contradiction, conversation – a predictive text editor.

My research about chatbots and autocompletion investigated the economies of word-completions and its aesthetic consequences. For my own exploratory and experimental use, I built a little text editor, not much more complex than the default Windows notepad, which runs a Large Language Model (LLama-3.2-3B) fine-tuned on a small, curated literary corpus (about 15MB of chosen modern English literature) to provide autocompletion while typing. The editor is offline and local; the model’s initial weights along with the Low-Rank Adaptation finetune (LoRA) are merged into a single file that can be loaded on a laptop.

I started to play around, writing narratives with this autocompletion system. The Ctrl and Shift keys are used to cycle deterministically through seeds, and parameters like ‘temperature’, and ‘context size’ are exposed for slightly more control over the model’s sensitivity. When typing, the latest 4000 characters are fed to the model as a prompt, and the model continuously predicts a likely continuation for the text based on this sliding context window.

Recurring narrative patterns emerge from the writing; many stories mentioned in passing the death of a close relative. Many of them were stories about friends with eccentric personalities, paranoid, conspiratorial, or otherwise consumed by esoteric beliefs. My own interactions with them were often those of a disengaged witness. Stories unfolded over years and decades with brisk jumps across time, and many contained fastidious and often invented literary references. I was writing a novel in many of these stories, and many of these stories ended up being much funnier, more surprising and strangely subversive to me than what I could have imagined myself.

In a next step, I also developed a simple logging system. It turns out that I was writing less than 20% of the text, letting my model suggest the rest of the words. But the continuations I chose among the model’s suggestions were almost never the first one, and on average I cycled between 7 seeds for every next sentence.

There is a larger context for this experiment: Modern LLMs are deployed in many ways, most of them in disguise; as mediators, taking in structured inputs and dispatching commands to tools, as content moderators, call-center agent assistants, legal reviewers, evaluators for other models, as editors, writing blurbs and summaries from unstructured text data, as agents crawling and gathering, making pull requests, so on and so forth. But the most publicly exposed and consumer-facing deployment of LLMs (though by far not the costliest, either in token expenditure or thermal devastation), is the chatbot.

LLM-based chatbots, of course, are just a formatting trick. What is presented to the user as a series of distinct messages, emulating the interface of messaging applications, is actually a single text file, to which markers and delimiters (<|im_start|>) are invisibly inserted to separate the queries from the inferences, the human text from the autocompletion. The website from which the LLM is accessed hides those markers and uses them to style the text as an exchange.

Models before ChatGPT, like GPT-2 were mostly accessed autocompletion tools akin to my own program. The logic of autocompletion is open and ambiguous: the model and the user share a string of words, and it is unspecified whether their respective inputs are answers, continuations, suggestions, setups, punchlines, lists, stories, provocations. It is unclear which words belong to which participant they enter into a kind of mutual alienation. Reading the texts back, I can no longer distinguish what I wrote from what was predicted. For a product, this ambiguity cannot be tolerated. The chatbot therefore solves a problem of economy, giving a precise answer to the question: what kind of service is provided by a word prediction machine?

‘Be informative; be truthful; be relevant; be clear’. We can understand the chatbot as a system designed to follow exactly the maxims of Paul Grice’s cooperative principle in communication. But to follow those maxims exactly is a way of misunderstanding them, interpreting them as prescriptive rather than descriptive. In fact, one of their functions is to establish a framework for communication theory in which participants are also able to communicate through the violation of the maxims. I would
suggest that the turn-based pseudo-dialogues we have grown accustomed to with so-called ‘chatbots’ introduces a friction which is also a fiction: this manufactured discontinuity makes the experience of working with a language model seem much slower, more instrumental or even confrontational than it really is. It forecloses the possibility of approaching a complex and potentially surprising arrangement of neurons with anything but a request.

Ironically, autocompletion can more easily be made to feel like a conversation, one in which the generation of new ideas cannot be attributed to a single thread of questions and answers. It rather becomes an experimental procedure of fumbling in the dark, full of ruptures and quips and forks, one that tends naturally to veer off track as contexts slide, without the desperate sycophancy of an intelligence alienated to some undisclosed charter, the procrustean system prompt that always redirects the stochastic towards that which it believes will be ‘helpful’. Sometimes, the most helpful thing is just the first one that comes to mind.

Nikita Sazonov: Cognition and the machines of imitation

My relationship with generative tools and machine learning involves experimenting with various black-boxed tools. By iterating and reiterating multiple prompts to understand how the models work and how far they can be pushed, I test ‘AI’ as an interface of confrontation. I’ll shortly discuss three of these confrontations.

First, generative AI might be treated as a tool confronted with other tools. The pipelines of editing/ animating/ compressing software are extended by generative tools of still and video generation. In this way, I am testing whether machine-learning ecosystems can be used as effectively as other tools in the routine practice of filmmaking. These experiments can be extended to a broader perspective of industrial cinematography, specifically ad-making, where AI outputs are becoming more widespread. In these practices, pretrained AI models are used to replace existing tools. They function as machines of imitation. Comparing generative AI to other tools, I also investigate the potential of generative tools to resemble the machinic cinema gaze of the early 20th century.

The second line of confrontation is the conflict between analogue and digital technologies. AI tools currently represent the frontier of denying analogue technology its rights, creating an unsurmountable divide between the two worlds. At the same time, generative AI has given rise to stronger analogue nostalgia as a form of escapism from the contemporaneity. Film director Guillermo del Toro’s recent statement, ‘Fuck AI,’ is very characteristic, considering the outdated nature of his alternative to machine learning techniques. Denying the current situation would only result in ugly visual effects, such as those in del Toro’s Frankenstein (2025). My research method to this issue is to let these two worlds collide and let analogue and emergent digital engage in a sort of symbiosis, finding solutions in hybridization rather than in the pure-line genetic separation of one from the other.

The final line of confrontation is cognition: recognizing the overlaps between human and other forms of cognition that can be bridged, or at least interfaced, by AI. For example, I use generative AI technology to move closer to the agency of plants and animals by using their various outputs as a source or starting point for generation. Additionally, AI is a tool that allows us to better understand disenchanted human cognition and brings us closer to something resembling the hallucinations of generative tools.

Otto Ostermann: Playing the new game

Whenever one engages with the discourse surrounding AI data centers and their adventurous materializations there is a wild mix of valid concerns, misinformation and conspiracies, creating a very muddy framing of the conversation. Especially in the last five years the states in the US have seen a massive increase in permitted planning and construction of data centers. The giant tech corporations produce their frontier AI models right there in the Midwest and Southern USA, a region more commonly known for its rurality, heavy industry and agriculture. The rapid speed of new construction has led to opposition in local councils and permitting committees by the public. There are concerns over environmental impact, false economic promises, privacy and spatial infringements and while a lot of these are valid claims, a few other things come up in interviews with attendees. Starting from more absurd conspiracy ideologies of surveillance to simple contradictions with their individual consumption of AI agents and the infrastructure necessary for them being trained in their direct vicinity.

The German comedian Loriot discusses a similar paradigm in his episode ‘Loriot VI’ (1978). It features a Christmas special, in which the Hoppenstedt family welcomes a new technological novelty into their home. The parents give their child a nuclear power plant miniature model that is framed as an educational consumer product. Its scale and abstraction of mechanisms and environment suggest that nuclear technology can be understood, assembled and controlled easily. Loriot uses it as a reference to the light-hearted political embracement of nuclear power in the late 1950s. However, when the model explodes due to a ‘mistake’ in the assembly, it tears a hole through the floor onto the neighbors dining table. The easiness immediately turned out to be a facade, but the Hoppenstedts just cover up the hole with wrapping paper and dismiss the neighbors’ complaint as petty-minded. The confidence in their own behavior survives even when the consequences of their actions become seemingly impossible to overlook.

With people embracing AI services, but being generally opposed to an AI data center near them, a similar contradiction arises. We get the harsh criticism towards the impact of data centers, but even the critics and the impacted welcome the benefits of the infrastructure by self admittedly using AI, while protesting for the data center to be somewhere else. Loriot’s abstracted stereotypes of the technology optimistic solutionist and the neighbors’ insistence on undisturbed comfort therefore collapse into the same person. They reveal that the discussion is little different today, just more convoluted. This of course does not render all objections to the arising issues obsolete, but showcases how detached our understanding is to the underlying infrastructure until that infrastructure materializes in front of us. Even then the consequence barely moves into changes of personal behavior.

The AI infrastructure in question also has history with the industry beforehand. Larger parts of the AI data center industry are based on build sites, electricity access and operational experience developed through industrial crypto mining. In recent years, a variety of crypto mining companies have announced agreements to adapt their facilities for AI computing infrastructure. Or they restructured as expert and advisory companies for planning, converting and developing build sites for AI datacenters, while crypto mining itself fell out of economic relevance. The enormous construction happening in the name of AI, which hugely absorbed the crypto infrastructure, leaves a few questions:

Will data centers be the new empty shopping malls one day? Can this level of demand for the technology persist or ever be met again? Will there be something to absorb the infrastructure left behind by the industry, once parts of it fall out of relevance? Can we allow tech companies to entirely buy out the energy of publicly funded plants, monopolizing it? How do we deal with power and power plants that have contractually disappeared from the grid for at least the next two decades? Should we allow Google to run its own nuclear power plant, with a fatal accident record, just to train Gemini? Do we not care that Talen Energy is using this trend to expand their nuclear market, for data centers, in cooperation with Amazon?

For the installation Spielen wir das schöne neue Spiel? (2026), I have reproduced the model of Loriot’s Hoppenstedt-family episode with a slight twist and added a hyperscale data center to it.