TLDR: Teams paired with artificial intelligence delegate more and edit less, flattening their output into sameness; new evidence shows varying the model’s persona restores distinctiveness, making homogenised work a configuration failure rather than an inevitable cost.
Human-AI teams ship more, and edit far less
The evidence on artificial intelligence in creative work has moved past anecdote and into controlled measurement, and the results are more interesting than either the enthusiasts or the sceptics predicted. Output rises. Quality rises on one axis and falls on another. Something else changes as well, quietly, and it takes a large sample to see it at all. A field experiment assigned 2,234 participants to human-human and human-AI teams and had them produce 11,024 advertisements for a think tank. The teams working with an artificial intelligence (AI) agent produced roughly 50 per cent more ads per worker, and their text scored higher on quality. All-human teams still made better images — a reminder that AI capability is jagged rather than uniform, strong on fluent language and weaker on composition, so the same tool changes a copywriter’s day and an art director’s day in opposite directions.
Then comes the finding that matters. The ads made by human-AI teams were measurably more similar to one another. The researchers call it diversity collapse: output that is individually good and collectively interchangeable. Each asset would pass a review on its own terms, and the portfolio as a whole loses the spread that distinguishes one brand’s voice from another’s. The effect stays invisible at the unit of work where most organisations look, because a creative director reviews one campaign against a brief and finds it strong, while nobody holds forty campaigns from four teams side by side and asks whether they have started to rhyme. Scale is what made the pattern legible, and the mechanism behind it turns out to sit in human behaviour rather than in the model.
The collapse comes from the handoff
What people do with an AI teammate differs systematically from what they do with a human one, and the difference is measurable in the interaction logs rather than inferred from the output. Participants delegated 17 per cent more work to AI agents than to human partners, and performed 62 per cent fewer direct edits on the text that came back. They also stopped talking like teammates: task-oriented messages rose by a quarter while interpersonal messages fell. The conversation narrowed to instruction and receipt, which is how a person talks to a service rather than how two people talk while building something together. Every one of those shifts points the same way, toward a working relationship in which the human sets the task and receives the artefact.
| Behaviour or outcome | Human-AI team vs. all-human team |
|---|---|
| Ads produced per worker | ~50% more |
| Work delegated to the partner | 17% more |
| Direct edits made to the partner’s text | 62% fewer |
| Task-oriented messages | 25% more |
| Interpersonal messages | 18% fewer |
| Text quality | Higher |
| Image quality | Lower |
| Similarity between finished ads | Higher (diversity collapse) |
Put plainly, people manage an AI agent instead of collaborating with it. They hand off, they accept, they move on. This is automation bias doing its familiar work — the long-documented human tendency to treat a machine’s confident output as sufficient and to stop scrutinising it. Fluency is the trigger: a draft that arrives grammatical, structured and plausible presents far fewer visible handholds for objection than a colleague’s rough paragraph, so the reviewer’s attention finds nothing to catch on and slides to approval. A rough draft invites intervention by looking unfinished, which is a property worth mourning now that most drafts arrive finished-looking whether or not they are any good.
Editing is where judgement, taste and context get pressed into the work, and that is precisely the step being skipped. A human editor rewriting a sentence imports everything the model never had: the client’s history, the sales meeting last Tuesday, the phrase a regulator objected to two years ago, the instinct that a particular metaphor will read as glib to a cardiologist. Remove the intervention and the output reverts toward the model’s central tendency, which is identical for every team using it. Delegation raised average text quality and, in the same motion, drove the convergence — the two results are one result, produced by a single behavioural change rather than by two independent properties of the tool.
One persona repeats itself; ten personas produce range
This is where the question turns from diagnosis to design. A separate study prompted generative artificial intelligence with ten distinct personas to generate 300 story plots, then measured how alike the results were using text embeddings — a numerical fingerprint of meaning, where a score near 1.0 means near-identical and a score near 0 means unrelated. Ideas drawn from any single persona were almost interchangeable, averaging 0.92 similarity. Ideas drawn across different personas averaged 0.20. Varying the persona at the input stage preserved the diversity of the collective output against a human-only baseline. The gap between those two figures carries the whole argument: the same model, asked the same underlying question, produced either near-duplicates or genuine range depending entirely on how the request was framed.
The mechanism is legible once stated. A generative model produces text conditioned on the context it is given, and a fixed prompt supplies a fixed context, so the distribution it samples from barely moves between requests. Changing the persona changes the conditioning, which relocates the model to a different region of what it can produce. An organisation that standardises on one house prompt has effectively chosen a single point in that space and instructed every team to draw from it, which explains why the outputs converge without anyone deciding that they should. The standardisation usually arrives for good reasons — consistency, quality control, speed of onboarding — and that is exactly why it goes unexamined.
The authors’ conclusion is the useful part: the trade-off emerges from uniform deployment practice rather than from a hard limit of the technology. Treat the model as a configurable partner and the range returns. Treat it as a single static voice that everyone queries the same way, and the organisation converges. So the honest answer to the question — is sameness fixable, or the price of the speed? — is that it is largely fixable, and most teams are simply paying the price by default. Default is the operative word, because nobody in a normal organisation ever proposed uniform output as a goal; it arrived as a side effect of decisions taken for other reasons entirely.
In regulated industries, sameness is a commercial risk
Commercial teams in pharmaceuticals, medical technology and financial services now run on the same handful of models, prompted in much the same way, inside review processes built to reward the safe and the familiar. That combination compounds in a way it does not in unregulated categories. The regulatory frame already compresses what a brand may say: promotional claims sit under Swissmedic and European Medicines Agency (EMA) scrutiny, device messaging under the EU Medical Device Regulation (MDR), and financial promotion under the Swiss Financial Market Supervisory Authority (FINMA). Add an AI layer that quietly pulls every team toward the median, and distinctiveness erodes from both ends at once — the permissible range narrows from outside while the tooling narrows the occupied space from within.
Review culture supplies the third pressure, and it is the one insiders underestimate because it feels like diligence. A medical, legal and regulatory board is designed to catch what is wrong, and a familiar-sounding claim resembling last year’s approved language passes with less friction than an unfamiliar one that is equally defensible. Reviewers therefore reward the median without intending to, since precedent is the cheapest available evidence that a phrasing is safe. An AI draft arrives pre-tuned to that same median, which means the tool and the reviewer now agree with each other for entirely different reasons. Three forces push in one direction, and none of them announces itself.
The damage accumulates slowly and stays easy to miss, because each individual asset clears review and reads well. The cost surfaces later, in campaigns nobody remembers and a brand that sounds like its competitors at the exact moment a customer is choosing between them. In categories where clinical differentiation between products is genuinely narrow, the brand’s distinctiveness carries a disproportionate share of the commercial work — it becomes the reason a prescriber recalls one name rather than another when two products carry comparable evidence. Erosion of that asset shows up as a slow drift in unaided recall and a gradual need to spend more for the same share of attention, both of which get attributed to media inflation or a tougher market rather than to a homogenised voice. That makes it an expensive thing to leave unmeasured, and expensive in a way no line item ever records.
Variance belongs in the dashboard alongside volume
The corrective is procedural, and it starts with measurement, because an organisation that tracks only throughput has no instrument capable of detecting the problem at all. For hiring and commercial leaders. Measure variance alongside volume: sample a quarter of AI-assisted output and test how similar it is to itself. Vary the input deliberately, with different personas, briefs and framings across teams rather than one house prompt for everyone. Rebuild the edit step into the workflow, because human judgement is the ingredient that disappears when an AI agent is treated as a subordinate. And hire for that judgement — the scarce profile is the brand-tech specialist who can direct a model, reject its first answer, and put a point of view back into the work.
Making the edit step survive contact with a deadline requires structure rather than encouragement. A team told to edit more will edit more for two weeks and then revert, because the draft still arrives fluent and the deadline still arrives on Friday. What holds is a workflow that separates generation from revision — a named owner for the revision pass, a rule that no first output ships, and a brief expectation that the reviewer records what they changed and why. That record doubles as evidence during review, which is the sort of dual benefit that makes a process stick in a regulated organisation, because a step serving two masters survives the next efficiency drive.
For professionals. The evidence points squarely at the skill worth building. Speed has become commodity; editorial judgement, taste and the confidence to override a fluent draft remain scarce. Keep your hands in the work — those 62 per cent of edits that colleagues now skip are exactly where a professional’s value becomes visible. A practitioner who can articulate why a competent draft is wrong for this audience holds a position that gets more valuable each time the tooling improves, because improving fluency raises the premium on the person willing to disagree with it. The next several years will sort teams by how many such people they managed to keep.
Edward Galle sources and develops brand-tech talent for regulated industries — the people who direct AI rather than defer to it. Employers can open a position oder review the services for companies. Professionals can submit a CV, and sharpen the judgement these roles demand through the Harvard Business Review training journeys.
References
- Collaborating with AI Agents: Field Experiments on Teamwork, Productivity, and Performance. arXiv. https://arxiv.org/abs/2503.18238
- Diverse AI Personas Can Mitigate the Homogenization Effect in Human-AI Collaborative Ideation. arXiv. https://arxiv.org/abs/2504.13868
- Diverse AI personas can mitigate the homogenization effect in human-AI collaborative ideation. ScienceDirect. https://www.sciencedirect.com/science/article/pii/S294988212600040X