Multimodal College-level Writing Dataset

A dataset of authentic student essays, detailed instructor feedback, and transcripts of one-on-one student–instructor conferences. This is built to model how writing actually develops, not just how it's corrected.

Student essays

Multiple drafts of the same essay, collected as it's revised across a semester.

Instructor feedback

Detailed written comments instructors leave directly on each draft — an average of 11 per essay.

Conferencing dialogue

Transcripts of one-on-one conversations where the student and instructor discuss the essay and its feedback.

Intended impact

  • Enable AI systems that model higher-order writing development, not just grammar.
  • Provide a public benchmark for evaluating and training educational AI feedback models.
  • Support equity research on how feedback effectiveness varies across ESL learners and demographics.
  • Bridge learning sciences and NLP communities through open, multimodal data.

Possible Use Cases

  • Automated essay scoringRubric-aligned, trait-specific assessment
  • Feedback generationBenchmarking AI vs. authentic instructor feedback
  • Revision outcome predictionPredicting draft improvement from feedback features
  • Idea development modelingTracking central idea evolution across drafts
  • Multimodal dialogue modelingTraining AI writing tutors on real conferencing data
  • Feedback quality evaluationClassifying feedback by pedagogical function
XL

Dr. Xiang Lorraine Li (PI)

Department of Computer Science

DL

Dr. Diane Litman (Co-PI)

Department of Computer Science / Learning Research & Development Center

GR

Dr. Gayle Rogers (Co-PI)

Department of English

RC

Raquel Coelho (Co-PI)

Department of Informatics & Networked Systems / Learning Research & Development Center