Impact of LLMs on Creativity

Short version: ==LLMs like ChatGPT can help produce creative looking ideas for one person, but when lots of people use the same model, the overall pool of ideas tends to get more similar.==[‌:cite[1]{ln=2}‌][‌:cite[2]...

Short version: ==LLMs like ChatGPT can help produce creative looking ideas for one person, but when lots of people use the same model, the overall pool of ideas tends to get more similar.==[‌:cite[1]{ln=2}‌][‌:cite[2]{ln=1}‌] What they seem to help with LLMs can sometimes match or beat humans on some standard creativity tests, including divergent thinking and remote association tasks.[‌:cite[3]{ln=1}‌][‌:cite[3]{ln=2}‌][‌:cite[3]{ln=3}‌] The paper also says emerging research suggests LLMs can at times outperform people on some established creativity measures.[‌:cite[4]{ln=2}‌] In this study’s setup, some prompting changes and parameter changes increased the individual diversity of GPT 4 essays, and one parameter modified GPT 4 condition even exceeded human written essays on the study’s individual diversity measure.[‌:cite[6]{ln=2}‌][‌:cite[6]{ln=3}‌][‌:cite[5]{ln=2}‌] Where the downside shows up The bigger problem is group level sameness.[‌:cite[7]{ln=1}‌][‌:cite[7]{ln=3}‌] Across three preregistered studies on 2,200 college admissions essays, each additional human written essay added more new ideas to the overall pool than each additional GPT 4 essay.[‌:cite[1]{ln=4}‌][‌:cite[1]{ln=5}‌][‌:cite[8]{ln=1}‌][‌:cite[8]{ln=2}‌] The authors found human writing increased collective semantic diversity about two to eight times more than base GPT 4 writing across the studies.[‌:cite[8]{ln=3}‌] In Study 2, for example, the diversity growth rate for base GPT 4 essays was only about 31% of the human rate.[‌:cite[9]{ln=4}‌][‌:cite[9]{ln=5}‌][‌:cite[9]{ln=6}‌] In Study 3, base GPT 4 was only about 11% of the human rate, and its increase was not statistically significant.[‌:cite[10]{ln=1}‌][‌:cite[10]{ln=2}‌][‌:cite[10]{ln=3}‌] So what’s the real takeaway? ==ChatGPT may boost or at least support individual creativity, but it can still shrink collective creativity by steering many users toward similar outputs.==[‌:cite[12]{ln=3}‌][‌:cite[11]{ln=2}‌] That pattern held even after attempts to make GPT 4 outputs more diverse through prompt engineering, parameter changes, and Chain of Thought prompting.[‌:cite[1]{ln=6}‌][‌:cite[8]{ln=4}‌][‌:cite[13]{ln=1}‌][‌:cite[13]{ln=4}‌] Even when GPT 4 became more diverse at the individual level, humans still produced higher group level diversity growth.[‌:cite[8]{ln=5}‌][‌:cite[14]{ln=5}‌][‌:cite[13]{ln=4}‌] Why this may happen The paper points to two likely reasons.[‌:cite[15]{ln=2}‌] First, LLMs are optimized to predict the most probable next token, which pushes them toward predictable and conventional outputs over surprising ones.[‌:cite[15]{ln=3}‌] Second, instruction tuning and reinforcement learning from human feedback can further encourage safe, aligned, and therefore more homogenized responses.[‌:cite[15]{ln=4}‌][‌:cite[15]{ln=5}‌] The authors also say LLMs are trained on existing data distributions, which can bias them toward predominant cultural ideas and away from unique perspectives.[‌:cite[12]{ln=1}‌][‌:cite[12]{ln=2}‌] Practical advice If you use ChatGPT for creativity, use it as a sparring partner, not as the source of truth.[‌:cite[16]{ln=2}‌][‌:cite[16]{ln=4}‌] It may help you get unstuck or increase idea fluency, but overreliance can flatten originality across teams, classrooms, or industries.[‌:cite[16]{ln=2}‌][‌:cite[16]{ln=3}‌][‌:cite[11]{ln=4}‌] The paper suggests monitoring for saturation and intervening with parameter changes, diversified data, better prompting, or human oversight to preserve variety.[‌:cite[17]{ln=2}‌][‌:cite[17]{ln=3}‌] Important limit This evidence is not a blanket statement about all creativity everywhere.[‌:cite[18]{ln=1}‌][‌:cite[18]{ln=3}‌] This paper only studied college admissions essays, and its main creativity measure focused on divergent idea diversity rather than things like logical coherence, practical usefulness, or real world problem solving quality.[‌:cite[18]{ln=2}‌][‌:cite[19]{ln=2}‌] So the safest conclusion is: ==LLMs can help individuals generate creative content, but widespread use may homogenize ideas at scale.==[‌:cite[11]{ln=2}‌][‌:cite[20]{ln=2}‌] If you want, I can also give you the plain English version for school/work , or the more technical research version .