Right now synthesia and hay gen are freaking out seeing this, if it's released they'll have to scrap everything.
Actually it's possible to create this, it's just the algorithm doesn't know how to think ahead, like:
NYC, person walking, populated scene, predict based on factors of matrix multiplication of vector, repeat x amount of times.
If you try that now it's possible but you'll get 1 frame per second, and controlling it will be difficult as it would be a dream scale of mad insanity while looking life life and incoherent and there's no deep neural network layer working on predictive behaviour other then then the next pixel and your input.
So it's possible just extremely time consuming.