For a project I'm working on, I require a function which copies the contents of a rectangular image into another via its pixel buffers. The function needs to account for edge collisions on the destination image as the two images are rarely going to be the same size.
I'm looking for tips on the most optimal way to do this, as the function I'm using can copy a 720x480 image into a 1920x955 image in just under 1.5ms. That's fine on its own, but hardly optimal.
#define coord(x, y) ((void *) (dest + 4 * ((y) * width + (x))))
#define scoord(x, y) ((void *) (src + 4 * ((y) * src_width + (x))))
void copy_buffer(uint8_t* dest, int width, int height, uint8_t* src, int src_width, int src_height, int x, int y) {
if (x + src_width < 0 || x >= width || y + src_height < 0 || y >= height || src_width <= 0 || src_height <= 0)
return;
for (int line = std::max(0, y); line < std::min(height, y + src_height); line++)
memcpy(coord(std::max(0, x), line), scoord(-1 * std::min(0, x), -1 * std::min(0, y)), (std::min(x + src_width, width) - std::max(0, x)) * 4);
}
Some things I've considered
Multithreading seems suboptimal for several reasons;
- Race conditions from simultaneous access to the same memory region,
- Overhead from spawning and managing separate threads
Using my system's GPU
- Effectively multithreading
- Huge overhead for moving and managing data between GPU and CPU
- Not portable to my target platform
Algorithmic optimisations such as calculating multi-image bounding boxes and adding **** loads more code to only render the regions of the image that will be visible
- While I was planning on doing this anyway, I thought I'd mention it here to ask for further information on how to best achieve this
Using a library/os function to do this for me
- I'm new-ish to programming on the low level, and especially to performance-oriented programming, so there's always the chance I've missed something.
- I'm presently not using a multimedia framework like SFML, because I'm trying to focus on executable and codebase size, but if that's the best idea, so be it.
Whew, bit of a mouthful. I apologise, but I would seriously appreciate any pointers.
Extra notes: I'm writing for/on Linux embedded devices over the DRI/M interface.
Edit
As per @Jérôme Richard's comment, some information about my system
Development machine: Dell inspiron 15 7570, 16GB RAM, i7 8core + Ubuntu 21.04 Target machine: Raspberry Pi 3B (1GB RAM, Broadcom something-or-other) 4 cores 1.4GHz + Ubuntu Server for Pi
Compiler: GCC/G++ 11.2.0