The students developed several validation approaches using large language models to check report completeness and compliance with predefined criteria, organised in a multi-agent architecture where agents cross-review each other's decisions. Experiments showed that this setup can improve the consistency and quality assurance of automated validation while keeping humans in the loop, pointing towards scalable AI-assisted quality control for real-world mystery shopping programmes.