Optimizing the Full Stack for Generative Image and Video Models┃VoiceTube - Learning English through Videos