DOI
10.31390/lsucontrol.26.5
COinS
Scene-Aware Multimodal Grounding for Verifiable Robot Task Planning