Quote of the day

A person who never made a mistake never tried anything new.

- Albert Einstein

Contact

Multimodal Rubric-Based Assessment of Handwritten Scientific Work with Calibrated Confidence

Open Education

Multimodal Rubric-Based Assessment of Handwritten Scientific Work with Calibrated Confidence

Multimodal Rubric-Based Assessment of Handwritten Scientific Work with Calibrated Confidence

Description

Assess handwritten scientific work containing equations and diagrams using explicit rubrics, while referring low-confidence cases to instructors.

Details

Context and Problem Statement

Grading handwritten scientific work is time-consuming and subject to inter-rater variability. Automated systems remain fragile on mathematical handwriting, where recognition errors can substantially alter interpreted reasoning.

Research Question

What level of reliability can multimodal rubric-based assessment achieve on handwritten scientific work, and what proportion of submissions must be referred to instructors to maintain acceptable fairness?

Proposed Approach

Combine mathematical handwriting recognition, reasoning-step segmentation, criterion-level scoring with a multimodal model, and confidence estimation that triggers human review. Compare model-human agreement with inter-grader agreement.

Expected Contribution

An evaluation framework that uses inter-rater disagreement as a realistic reference and a selective-review strategy based on uncertainty.

Expected Prototype

An assisted-grading application providing criterion-level scores, linked explanations, and a queue of uncertain submissions for instructor review.

Datasets

CROHME and anonymized double-graded student work collected under institutional approval.

Challenges

Handwriting variability, recognition-induced grading errors, presentation bias, legal constraints, and stakeholder acceptance.

Research Question

Can multimodal rubric-based grading achieve agreement with human graders comparable to inter-grader agreement when uncertain cases are selectively referred for review?

Innovation

The project uses human inter-rater disagreement as the realistic reference and designs the system around selective escalation rather than maximizing automation.

Expected Deliverable

An assisted-grading prototype, a double-graded corpus, and a comparative analysis of human-machine agreement.

Technologies

Multimodal LLMs and Document Vision Mathematical Handwriting Recognition Rubric-Based Assessment Uncertainty Quantification Human-in-the-Loop

Required Skills

  • Computer Vision and Multimodal Models
  • Natural Language Processing and Mathematical Reasoning
  • Inter-Rater Agreement Statistics
  • Ethics of Educational Assessment

Datasets

  • CROHME
  • Anonymized double-graded student submissions
  • Rubrics designed with instructors

Morocco & Africa Relevance

Large class sizes in Moroccan universities and engineering schools make detailed formative feedback difficult; reliable assistive grading can reduce workload while preserving human oversight.