Multimodal Rubric-Based Assessment of Handwritten Scientific Work with Calibrated Confidence
Multimodal Rubric-Based Assessment of Handwritten Scientific Work with Calibrated Confidence
Description
Details
Context and Problem Statement
Grading handwritten scientific work is time-consuming and subject to inter-rater variability. Automated systems remain fragile on mathematical handwriting, where recognition errors can substantially alter interpreted reasoning.
Research Question
What level of reliability can multimodal rubric-based assessment achieve on handwritten scientific work, and what proportion of submissions must be referred to instructors to maintain acceptable fairness?
Proposed Approach
Combine mathematical handwriting recognition, reasoning-step segmentation, criterion-level scoring with a multimodal model, and confidence estimation that triggers human review. Compare model-human agreement with inter-grader agreement.
Expected Contribution
An evaluation framework that uses inter-rater disagreement as a realistic reference and a selective-review strategy based on uncertainty.
Expected Prototype
An assisted-grading application providing criterion-level scores, linked explanations, and a queue of uncertain submissions for instructor review.
Datasets
CROHME and anonymized double-graded student work collected under institutional approval.
Challenges
Handwriting variability, recognition-induced grading errors, presentation bias, legal constraints, and stakeholder acceptance.
Research Question
Innovation
Expected Deliverable
Technologies
Required Skills
- Computer Vision and Multimodal Models
- Natural Language Processing and Mathematical Reasoning
- Inter-Rater Agreement Statistics
- Ethics of Educational Assessment
Datasets
- CROHME
- Anonymized double-graded student submissions
- Rubrics designed with instructors