I enjoy playing with AI-based image generators, which are a source of constant surprises. Sometimes someone will come up with a clever concept and briefly everyone is enchanted with it. For me, the enchantement is “how did they do that?” and the subsequent pleasure of finding out.
This one is a wowser. Here’s how the trick is done. First off, I created a plausible base image to work from. The prompt, in this case, was:
a detailed cinematic image of a samurai walking down a misty trail, wearing a conical straw hat, he has his katana out of the scabbard. behind him on the trail are several crumpled bodies. one is still alive, and has an expression of pain and horror in its face. he has his back to the trail and is walking toward the camera.
detailed, realistic. black and white in the style of akira kurosawa.

Well, the black and white version is better, but no biggie. Load it into photoshop and draw a camera-track on it:

That’s basically all the set-up. The rest is pure fun. Now, the prompt for the video, run through seedance2.5 to render it, results in:
Right, what does the prompt do? It tells the AI to follow the track that I drew on the image.
Image-to-video, 30 seconds, ONE continuous flying camera shot, no cuts. FIRST FRAME: the video starts as an EXACT copy of @Image1 —a frame from a black and white samurai movie, includingthe red drawn line and numbers “1” “2” and “3”. Within the first half second the red markings dissolve completely and must NEVER reappear. CONCEPT — FROZEN TIME: as the markings vanish, the photograph gains full three-dimensional depth, but TIME STAYS COMPLETELY FROZEN. Every person,vehicle, bird and particle is locked mid-motion like a vast sculpture: swirling mist hanging solid in the air,walking man frozen in mid-step. NOTHING moves except the camera, which drifts through the frozen moment slowly anddeliberately, inspecting details and faces at close range.
Basically, the camera is invited to confabulate detail into the image, to let it fly around and see things it otherwise would not see. But the image reference powerfully anchors it to the original scene.
KEEP THE ORIGINAL BLACK-AND-WHITE LOOK: monochrome tonality, fine photograin, overcast 1930s daylight — a living archival photograph. CAMERA PATH: the red line is the flight trajectory — enter at marker “1”, go to marker “2” and end at marker “3” . 00–01s: Static frame identical to the input; red markings dissolve. Dustmotes hang motionless in the air — the first hint that time is frozen. 01–04s: The camera dives to marker 1: macro pass over the scared man, the mist in the air is frozen like stone 04–08s: The camera rises to the samurai’s face frozen mid-motion: extreme close-up — stubble, creased eyes, a wisp of frozen breath vapor at his lips. Slow half-orbit around his head, focus racking across his face. 08–12s: Glide to the the dead and wounded on the trail. Close orbit — their clothes, their slack expression, unfocused eyes. 12–16s: The camera sweeps down the misty trail, swirling around the lens in sharp macro focus.
There are some really cool versions of this on instagram, done by people who spent the money for the high-end rendering pipelines. Seedance isn’t free, either, but it’s not particularly expensive. Some of the instagram videos are elaborate live-action scenes derived from famous paintings (e.g.: the girl with the pearl earring talking to Vermeer) and the woman with the ermine nose-booping it as she gets it out of its cage.
All of this is going to contribute to additional moral panics, when people realize that “OMG someone could make fake video!” Well, yes, and it’s going to be obviously fake, if only because it looks too good.

Sorry, but I can’t get past the basics that you specifically asked for B&W and got color, and that you specified that the wounded man should be behind the samurai, but he winds up in the foreground. The fact that you were able to make something you found useful or entertaining from what it gave you is beside the point. It didn’t follow basic instructions for step one.