Few-Shot Vision-Language Inspection with Verifiable Explanations on Production Lines
Few-Shot Vision-Language Inspection with Verifiable Explanations on Production Lines
Description
Details
Context and Problem Statement
Industrial quality inspection suffers from extreme defect scarcity and frequent product changes. Vision-language models can support few-shot inspection, but their textual explanations may not correspond to the actual defect location, reducing operational trust.
Research Question
How can consistency among classification decision, defect localization, and textual explanation be aligned and verified in a few-shot vision-language inspection system?
Proposed Approach
Adapt vision-language models through prompt learning and normal-only examples, add a localization head, verify agreement between anomaly maps and generated descriptions, and distill the system into a compact model for real-time deployment.
Expected Contribution
A quantitative protocol for evaluating explanation faithfulness in industrial anomaly detection and the performance cost of enforcing consistency.
Expected Prototype
A real-time inspection station that highlights suspicious regions, produces corresponding explanations and confidence scores, and supports rapid addition of new defect types.
Datasets
MVTec AD, VisA, and partner-line images annotated by experts.
Challenges
Extreme class imbalance, lighting variability, latency constraints, expert-annotation cost, and plausible but unfaithful explanations.
Research Question
Innovation
Expected Deliverable
Technologies
Required Skills
- Computer Vision and Multimodal Models
- Real-Time Optimization and Deployment
- Industrial Vision and Lighting
- Experimental Evaluation Methodology
Datasets
- MVTec AD
- VisA (Visual Anomaly dataset)
- Images collected from a partner production line