MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education · Bharat Hunt