Researchers from Iqra University and Ulster University have developed a temporally aware, intra-video Retrieval-Augmented Generation (RAG) framework to improve VideoQA accuracy for lecture videos. This framework aligns speech transcripts and visual captions to temporal boundaries, and refines retrieved segments with a cross-encoder before answer generation. The method was evaluated on the LectQA-Vid dataset, demonstrating improved factual alignment and robustness over non-temporal baselines.