
Concerns are growing among researchers about the capabilities of artificial intelligence tools like ChatGPT, particularly in how they evaluate and generate creative text.
A recent study highlighted that popular AI language models — including ChatGPT — can be easily misled into rating nonsensical or pseudo‑literary text as high‑quality. This finding suggests that, despite advances in natural language processing, AI systems might struggle to distinguish between genuinely meaningful writing and cleverly‑constructed gibberish.
The research, presented by a German academic, involved feeding the AI increasingly convoluted sentences and asking it to rate them for literary merit. Surprisingly, even clearly absurd or “made‑up” sentences were often scored highly by the model, even with its reasoning‑enhancement features activated.
Experts say the finding underscores broader challenges in the development of generative AI — particularly how these systems are trained to mimic human‑like judgments. Unlike human readers, AI models rely on patterns in data rather than true comprehension, which can lead them to mistakenly valorize stylistic oddities or incoherent text.
Critics argue that this limitation matters not just for literary applications but also for broader information quality. Other studies have shown that language models can produce convincing‑sounding but incorrect or fabricated content — a phenomenon often called “hallucination” or “confabulation” in AI research.
The debate reflects wider questions about how AI systems should balance fluency and accuracy, and whether current models can ever fully replicate the nuance of human judgment when it comes to creative expression or evaluating factual content.