It is possible, and relatively simple to seek and extract frames without loading whole movie into memory.
There is some speed-up loss, when the requested frame to seek is not a key frame.
When FFmpeg is requested to seek a non-key frame, it seeks to the closest key frame before the requested frame, and decodes all the frames from the key frame to the requested frame.
The demonstrating code sample does the following:
- Builds synthetic 1fps video with a running frame counter - perfect for testing.
- Execute FFmpeg as sub-process with stdout as an output PIPE.
The code sample seeks to the 11'th second, and set duration to 5 seconds.
- Read (and display) decoded video frames from the PIPE until there are no more frames to read.
Here is the code sample:
import numpy as np
import cv2
import subprocess as sp
import shlex
# Build synthetic 1fps video (with a frame counter):
# Set GOP size to 20 frames (place key frame every 20 frames - for testing).
#########################################################################
W, H = 320, 240 # video dimensions
video_path = 'video.mp4' # path to video
sp.run(shlex.split(f'ffmpeg -y -f lavfi -i testsrc=size={W}x{H}:rate=1 -vcodec libx264 -g 20 -crf 17 -pix_fmt yuv420p -t 60 {video_path}'))
#########################################################################
# ffmpeg command
command = [ 'ffmpeg',
'-ss', '00:00:11', # Seek to 11'th second.
'-i', video_path,
'-pix_fmt', 'bgr24', # brg24 for matching OpenCV
'-f', 'rawvideo',
'-t', '5', # Play 5 seconds long
'pipe:' ]
# Execute FFmpeg as sub-process with stdout as a pipe
process = sp.Popen(command, stdout=sp.PIPE, bufsize=10**8)
# Load individual frames in a loop
nb_img = H*W*3 # H * W * 3 channels * 1-byte/channel
# Read decoded video frames from the PIPE until no more frames to read
while True:
# Read decoded video frame (in raw video format) from stdout process.
buffer = process.stdout.read(W*H*3)
# Break the loop if buffer length is not W*H*3 (when FFmpeg streaming ends).
if len(buffer) != W*H*3:
break
img = np.frombuffer(buffer, np.uint8).reshape(H, W, 3)
cv2.imshow('img', img) # Show the image for testing
cv2.waitKey(1000)
process.stdout.close()
process.wait()
cv2.destroyAllWindows()
Note:
The argument -t 5 is relevant when the playing duration is known from advance.
In case the play duration is unknown from advance, you can remove the -t and break the loop when needed.
Time measurements:
- Measure reading all frames at once.
- Measure reading frame by frame in a loop.
# 6000 frames:
sp.run(shlex.split(f'ffmpeg -y -f lavfi -i testsrc=size={W}x{H}:rate=1 -vcodec libx264 -g 20 -crf 17 -pix_fmt yuv420p -t 6000 {video_path}'))
# ffmpeg command
command = [ 'ffmpeg',
'-ss', '00:00:11', # Seek to 11'th second.
'-i', video_path,
'-pix_fmt', 'bgr24', # brg24 for matching OpenCV
'-f', 'rawvideo',
'-t', '5000', # Play 5000 seconds long (5000 frames).
'pipe:' ]
# Load all frames into numpy array
################################################################################
t = time.time()
# run ffmpeg and load all frames into numpy array (num_frames, H, W, 3)
process = sp.run(command, stdout=sp.PIPE, bufsize=10**8)
video = np.frombuffer(process.stdout, dtype=np.uint8).reshape(-1, H, W, 3)
elapsed1 = time.time() - t
################################################################################
# Load load individual frames in a loop
################################################################################
t = time.time()
# Execute FFmpeg as sub-process with stdout as a pipe
process = sp.Popen(command, stdout=sp.PIPE, bufsize=10**8)
# Read decoded video frames from the PIPE until no more frames to read
while True:
# Read decoded video frame (in raw video format) from stdout process.
buffer = process.stdout.read(W*H*3)
# Break the loop if buffer length is not W*H*3 (when FFmpeg streaming ends).
if len(buffer) != W*H*3:
break
img = np.frombuffer(buffer, np.uint8).reshape(H, W, 3)
elapsed2 = time.time() - t
process.wait()
################################################################################
print(f'Read all frames at once elapsed time: {elapsed1}')
print(f'Read frame by frame elapsed time: {elapsed2}')
Result:
Read all frames at once elapsed time: 7.371837854385376
Read frame by frame elapsed time: 10.089557886123657
The result shows that there is some overhead of reading frame by frame.
- The overhead is relatively small.
There is a chance that the overhead is related to Python and not to FFmpeg.