Many tasks performed by language models have just one right answer. Code either compiles or it doesn't. A summary either reflects the source or misses it. However, creative writing breaks that pattern, as there is no ground truth for whether a story is good, but rather only readers who may disagree, sometimes sharply, about the same paragraph. That makes creativity one of the few remaining places where training and evaluating a model still depends on a person with taste, not a scoring function.
It also makes creativity a strange thing to train on at scale. Models trained on more AI-written text tend to produce writing that looks more like other AI-written text. As more of the writing online is AI-assisted, a clean, well-labeled sample of purely human creative output, created and scored by real writers and readers, becomes harder to find and more valuable to have.
We built a dataset to look at where machine judgment and human judgment on creative writing line up, and where they split. Sigma published this Creative Writing dataset for free on Hugging Face: 110 human-written story segments across 22 prompts, each response is scored by human reviewers as well as four LLM judges working from the same rubric, and a set of traditional NLP metrics is also computed for comparison.
This dataset could be useful in many ways – here are just a few:
-
Per-writer style analysis. Each writer answered up to 22 prompts under a stable pseudonym, enough responses per person to study how someone's vocabulary, rhythm, or scores hold steady, or don't, across different prompts.
-
Automated metric development. Lexical diversity, rare-word weighting, and sentence-rhythm shifts are calculated for comparison with human ground-truth ratings in the same rows. That pairing is what building or calibrating a creativity-scoring metric actually requires: a way to check whether the metric's number and a trained reader's judgment agree.
-
LLM-judge agreement and bias analysis. Four judge models scored the same responses on the same rubric, laid out side by side with human scores. In our own validation, Gemini scored consistently stricter than GPT-4o, and stricter than the human reviewers. That's not a footnote. It means the choice of judge model changes the verdict, and anyone using an LLM to grade creative work should know that before they build on it.
-
Human inter-rater reliability. Most responses carry more than one independent human reviewer, so the disagreement between people, not just between people and machines, is visible in the data too.
None of this settles whether AI narrows creative output or expands it. What it does is give researchers, and anyone building evaluation pipelines for creative work, a dataset where the writing, the AI scores, and the human scores sit in the same row, ready to be checked against each other rather than taken on faith.
That's the resource getting harder to find: writing untouched by AI, scored by people who can tell good from merely fluent. Sigma builds it any way, in a form researchers can put to the test themselves.
Explore the dataset on Hugging Face.
