Does AI already have a taste of its own?

7 min read

Generated images seem infinite, yet many converge on the same idea of beauty. The real creative challenge is not obtaining a good image, but building a language that does not feel pre-decided by the machine.

Open an image generator and ask for something simple: a city of the future, the portrait of a writer, a brutalist interior, a cinematic scene in the rain. Within seconds, images appear that differ in subject and detail yet often look surprisingly similar in the way they try to persuade us. The light is perfect, the contrast decisive, the composition legible, the surfaces polished. Everything seems ready to become a cover, a poster or a film still.

The sensation is familiar: we face almost infinite possibilities and still recognize the same taste. It is not style in the fullest sense, because it does not emerge from a biography, a conflict or an expressive necessity. It is closer to an aesthetic average: the meeting point between images in the data, the criteria used to optimize models and the preferences that millions of users continue to reward.

Artificial intelligence has no personal taste. It does not love one kind of light more than another and has no memories attached to a color. But it knows which configurations are most often associated with words such as “beautiful,” “cinematic,” “elegant” or “powerful.” When a request remains generic, the system tends to occupy that probabilistic territory, producing a recognizable, effective and culturally validated solution.

The problem, then, is not that AI images are ugly. Many are technically impressive and in several experiments are judged as pleasing as—or more pleasing than—images made by humans. The problem begins when pleasantness becomes automatic. An image can be seductive without containing a decision; it can look complete before it has found a reason to exist.

Convergence is not limited to color and lighting. It affects bodies, faces, environments and social roles. A study published in Scientific Reports found gender stereotypes and racial homogenization in faces produced by Stable Diffusion XL: different people were reduced to overly similar physical features, clothes and cultural signs. When a model compresses diversity into a few recognizable markers, it is not merely repeating an aesthetic. It is narrowing how we imagine the world.

Systems designed to “improve” prompts automatically also expose this mechanism. Research projects such as BeautifulPrompt train models to turn a simple request into a description capable of producing images rated as more beautiful. The result is useful, but raises a question: who defines that beauty? If an aesthetic score consistently rewards clarity, detail, harmony and spectacle, optimization may eliminate precisely the ambiguous, spare, uncomfortable or imperfect images from which a new language often begins.

A prompt, therefore, is not the same as a style. It can describe a palette, a lens, a material, an era or an art movement. It can impose rules and forbid clichés. But style is not a list of attributes. It is a coherent relationship between what an author sees, what the author excludes and how those choices change over time.

Cinema makes this obvious. A film’s cinematography is not recognizable because every shot contains the same keywords. It is recognizable because light, lenses, distance from bodies, movement, production design and color answer the same narrative intention. One scene may contradict the previous one and still belong to the same film. Coherence is not repetition; it is continuity of thought.

This continuity is more difficult with generative images. Every new request reopens the field of possibilities, and the model may reinterpret what appeared settled. The same sentence does not guarantee the same tone; a visual reference can preserve the surface while losing the structure; a palette can remain identical while the relationship between figure and space changes completely. Research on style transfer shows how difficult it remains to separate and control content, structure and aesthetic qualities.

The author’s work therefore moves from finding the perfect image to building a grammar. A grammar establishes relationships: how much empty space to leave, how to treat faces, which colors not to use, how explicit a symbol may become, where to accept error and when one image should remain quieter so that the next can emerge.

1. Define relationships, not only adjectives. “Cinematic” and “editorial” are too broad. It is more useful to specify how the figure inhabits space, where light comes from, which element should dominate and which should remain unresolved.

2. Build an archive of decisions. Saving successful prompts is not enough. Approved images, rejected images and the reasons behind each decision should be recorded. Style also emerges from the continuity of refusal.

3. Protect certain imperfections. Asymmetry, abrupt cropping, empty space and unpolished materials can prevent the system from pulling everything toward the most spectacular solution. Imperfection is not a filter added at the end; it must enter the image’s structure.

4. Judge the series, not the single result. A cover may work in isolation and weaken the identity of an entire project. Placing images side by side reveals repetitions, deviations and inconsistencies that remain invisible when each file is judged alone.

5. Change the model without losing direction. A genuinely defined visual language should survive, at least in part, a move from one tool to another. If it depends entirely on one model’s default aesthetic, it belongs more to the platform than to the author.

This does not mean locking oneself inside a formula. An overly rigid visual identity soon becomes another automation. Grammar exists to make evolution legible, not to prevent change. A coherent series must also be able to introduce a rupture, but that rupture should be a decision rather than a lapse of memory.

Research by Doshi and Hauser on AI-assisted writing revealed a paradox that is equally useful for images: the perceived quality of individual work can rise while overall diversity falls. Everyone gets a better result, but the results begin to resemble one another. In visual culture this effect can grow even stronger because the most spectacular images circulate, are imitated and return to the system as new references.

A loop forms: models propose what is most recognizable, users select what appears most successful, platforms display what receives the most attention, and that same aesthetic becomes the starting point for subsequent generations. Average taste is not imposed by an isolated machine. It is built collectively by data, models, interfaces, metrics and human behavior.

Leaving this loop does not require rejecting artificial intelligence. It means using AI against its own tendency to converge: asking for alternatives that are not merely cosmetic variations, comparing different tools, introducing personal material, deliberately limiting possibilities and above all recognizing when a good image is not yet our image.

AI does not possess taste in the way a person does. But it has enormous aesthetic force: it makes some solutions easier, faster and more available than others. Creative work begins by seeing that force while it operates. When everything looks immediately beautiful, authorship starts with asking who has already decided what beauty means.

  • AI aesthetics
  • Generative images
  • Visual culture
  • Prompting
  • Style
  • Authorship
  • Art direction
  • Homogenization
  1. Doshi and Hauser — Generative AI enhances individual creativity but reduces the collective diversity of novel content
  2. AlDahoul, Rahwan and Zaki — AI-generated faces influence gender stereotypes and racial homogenization
  3. Grassini and Koivisto — Understanding how people evaluate AI-generated artworks
  4. Cao et al. — BeautifulPrompt: automatic prompt engineering for text-to-image synthesis
  5. Wang et al. — InstantStyle-Plus: style transfer with content preservation