Why LLMs Give You Generic Creative Advice
An LLM hands you beige creative advice by design. Here is why that happens, the research that proves it, and how to use one without losing your edge.
I work at Moonb, and I use an LLM most days. So does most of my team. Which is exactly why I get irritated when someone shows me a chatbot transcript and asks why the “creative direction” it gave them reads like a horoscope written by a committee.
The honest answer is that the model did its job. You cannot prompt your way out of generic. Generic is the output the machine was built to produce, and once you understand why, you stop expecting originality from it and start using it for the thing it is actually good at.
The generic advice is the model doing its job
Ask any of the big models for a campaign idea and you get the same shape of answer. Lead with emotion. Tell a story. Show, do not tell. Make it authentic. None of it is wrong. All of it is useless, because it is the advice that applies to every brand, which means it applies to yours no more than to your competitor’s.
This is not a coincidence and it is not a bug someone will patch. A language model predicts the most probable next token given everything it has seen. “Most probable” is the operative phrase. It is engineered to land in the center of the distribution, and the center is where the safe and the already-said live. When you ask it to be creative, you are asking a system optimized for the expected to produce the unexpected. It will comply in the most expected way it can.
I find the “blurry JPEG of the web” description useful here. The model has compressed an enormous amount of human writing into a set of weights, and what survives compression is the average. The sharp, weird, specific stuff (the exact thing that makes creative work land) is the detail that gets smoothed away.
Why every model hands you the same beige middle
Three forces push the output toward the middle at once.
The training data is averaged. The model learned from billions of documents and its instinct is the consensus of all of them. Consensus is the enemy of distinctive creative.
The decoding anchors toward the common answer. Even with temperature turned up, the model gravitates to high-probability continuations. You can nudge it. You cannot relocate its center of gravity.
The safety and helpfulness tuning sands off the edges. Models are trained to be agreeable and inoffensive, and the sharpest creative ideas are frequently neither. The tuning that makes a model pleasant to talk to also makes it reluctant to say anything a focus group might flinch at.
Here is the part people miss when they think switching tools will save them. The homogenization is not confined to one model. Researchers have named the “Artificial Hivemind” effect, where different LLMs, trained by different labs, converge on strikingly similar answers to the same prompt. Jumping from one model to another does not restore originality. You are choosing between products that all learned from overlapping slices of the same internet.
And the effect compounds across people. In a set of studies on how models shape human expression, human-written essays produced roughly two to eight times more collective semantic diversity than base GPT-4 essays, and the flattening held even after researchers modified prompts and parameters specifically to boost variety. A separate comparison ran 22 different models against 102 people. The models were dramatically less diverse, with the gap widening as more outputs were generated. The more a room full of marketers all reach for the same tool, the more their work drifts toward the same place.

The finding nobody wants to hear
There is a study I keep sending to people, because it settles an argument I have had a hundred times. Published in Management Science, it tested how LLMs affect the quality of ad copy depending on who was using the model and how. Quality was measured by real clicks on live social platforms, not a survey.
The result splits cleanly. When the model was used as a sounding board, giving feedback on a human’s draft, it improved the work of non-experts. When it was used as a ghostwriter, generating the copy itself, the picture changed. For skilled creatives, letting the model write produced no benefit and was actively detrimental to quality, through an anchoring effect. The moment an experienced person saw the model’s version, their own thinking snapped to it. The draft became the anchor, and their expertise stopped fully firing.
Read that twice if you are a good creative who has been letting the model write your first drafts. The evidence says that is the one workflow that makes your output worse. Not neutral. Worse.
It maps to what I watch happen in the room. A strong art director stares at a model’s opening line and, instead of blowing past it, starts editing it. The ceiling on the work drops to whatever the model produced, because they anchored to it before their own instinct got a shot.
What actually gets eroded is your distinctiveness
None of this would matter much if distinctiveness were a nice-to-have. It is most of the job.
Nielsen puts it plainly: about 47% of an ad’s sales impact comes from the creative content itself, more than targeting or media placement. The creative is roughly half of what makes advertising work. So the exact quality an LLM sands away, the specific and unexpected stuff that reads as distinctly you, is the quality carrying the commercial result.
Now stack the adoption numbers on top. Around 91% of marketing teams now use AI in their work, and roughly 80% use it for content creation. Content creation is the single most common use case. So most of your category is pointing the same homogenizing tool at the same task, drifting toward the same middle, at the same time. If you let the model write, you are not just accepting average work. You are converging on the identical average as everyone else, in the one area that drives the sale. Distinctiveness stops being a creative preference and becomes the hedge.
How I actually use an LLM on real creative work
So we still use it. Every day. The trick is the frame: a brilliant sparring partner, a terrible ghostwriter. It is there to make my thinking sharper and more pressure-tested, never to hand me the finished thought.

Here is the split I hold to.
| Ghostwriter (avoid) | Sparring partner (use) |
|---|---|
| "Write the campaign concept." | "Here is my concept. Argue against it." |
| You anchor to its first draft | You write first, it stress-tests |
| Output drifts to the category average | Output stays yours, just harder |
| Worse for experts (per the research) | Helps non-experts, safe for experts |
The prompting tactics people trade around are real, but they only matter once you have accepted the model will never be the source of the idea. They make it a better sparring partner, not a better ghostwriter.
| Tactic | What it is actually good for |
|---|---|
| Persona ("react as a skeptical CFO") | Pressure-testing an idea from an angle you would not reach for |
| Constraints ("no adjectives, six words") | Forcing it off the obvious answer so it stops repeating itself |
| Chain-of-thought ("show your reasoning") | Seeing the logic so you can disagree with the step, not the output |
| Few-shot (paste your real brand voice) | Getting notes in your register instead of the generic one |
I lean on it hardest for divergence. Give me forty bad angles on this brief in ninety seconds. Most are useless, a few are wrong in a way that jostles something loose, and one occasionally reframes the problem. That is a fair trade, because I am the one deciding which of the forty is worth anything. The model widened the search. It did not make the call. If you want a deeper version of the input side, the way you frame the brief still decides most of the output, and I wrote about that in how to write a creative brief.
The workflow that keeps the taste in human hands
When we bring AI into real client work, it sits inside four stages, and a human owns the two that matter.
A human sets the brief. The problem, the tension, the audience, the thing that makes this brand not-that-brand. This is judgment, and the model has none of it.
The AI diverges. Volume, angles, provocations, first-pass variations. This is where quantity is a feature and where the model earns its place. Let it be prolific and slightly unhinged here.
A human curates and synthesizes. Someone with taste kills 90% of it, notices the odd fragment worth chasing, then combines pieces the model would never have joined. This is the stage the research is really about. It is where anchoring will hurt you if you skip it and just take the model’s favorite.
A hybrid pass executes. The human writes the actual thing, using the model as an editor and a checker, not an author. It can catch a weak transition or flag a cliche. It does not get to make the final line.
Notice what the human keeps: the brief and the curation. The two points where taste lives. The AI gets the two points where volume helps and taste is not yet required. That division is the whole game, and it is roughly how we think about where AI fits across creative work more broadly in will AI replace animators.
This is also why an embedded creative team still matters when everyone has the same models. The tools are commodity now. The judgment about which of the forty angles is the one, and the nerve to keep a brand distinct while the category converges on beige, is the part that does not compress into weights. If you want that judgment sitting inside your team instead of rented by the project, that is the gap Moonb was built to fill.
Use the model. Just do not let it hold the pen.
Frequently asked questions
Better prompts improve the output at the margin, so they are worth learning. But the homogenizing effect persists across prompt and parameter tweaks, which is why you should treat the model as a sparring partner rather than a source of original ideas. Use it to diverge and pressure-test, then make the creative call yourself.
No. The Artificial Hivemind effect means different models, trained by different labs, converge on very similar answers to the same query, so switching tools does not restore originality. The lever is how you use the model, not which one you pick.
The Management Science research says no. Letting the model ghostwrite is detrimental to expert quality through an anchoring effect, where your own thinking snaps to the model's version. Write the draft yourself and use the model to challenge and refine it instead.