TL:DR version of the question: my PRU sometimes lags its timings and throws off the devices downstream. How do I figure out why?
I'm using am AM335x CPU with two on-chip PRUs, to drive 32 strips of RGBW LEDs. A program running on the ARM CPU prepares the frame, and dumps it into PRU RAM. The PRU idles until the CPU sets a flag, at which point the PRU grabs the frame data 32 bytes at a time, and bitbangs them out.
Now, clearly, we're using the PRU because it can guarantee timings. I've gotta toggle high/low on timings which are the order of nanoseconds. And it works, at least most of the time. But there are points where the timing lags, and that throws off the data going down the wire and causes a flicker in the LEDs. It's definitely problems in timing- I've hooked up a logic analyzer and can spot situations where it spends 1083ns doing something it should have done in 600ns.
Here's the core code in clpru assembly. There's an additional include file which just handles mapping GPIO pins to register addresses, and assigns a few registers aliases, it contains no logic.
The actual logic for banging out starts here, where we set a register to point at the start of the frame and then read the current cycle counter into a register. (I'm using a fair number of aliases and macros just to try and keep this readable).
We then grab 32-bytes and stuff it into registers (scratch is register 11, so the 32-bytes spreads across r11-r18).
Then we take the first bit out of each byte and align it with an output pin. This is a big block of operations which takes up to 320ns to complete (two instructions in the macro, 32 executions of said macro).
Once everything is prepped we wait until we're 550ns from the last READ_TIME sleeper call. WAIT_NS reads the cycle counter in a busy loop.
The pattern for sending out for our LEDs is to send HIGH. For a zero, we hold approx 250ns (300ns is spec). For a one, we hold for approx 600ns. Then we hold until approx 1200ns have passed total.
We do this all relative to that BASE_WAIT of 550ns from before. At the end of the BIT_LOOP we reset the sleeper.
Again, 99% of the time (more!) this works fine. But once every few hundred frames, the timings get off. Something takes too long, and I can't understand why. And I definitely don't understand why it's fine almost all the time, and then jitters suddenly.
I have noticed that if a jitter happens, it almost certainly throws off the timing on fetching the next byte from PRU RAM. According to the docs, accessing PRU DRAM is deterministic. At the top and bottom of every frame, I do access DDR, just to check flags shared with the ARM program.
So my question is less "what's wrong" and more "how do I figure out what's wrong" or "what don't I understand about the PRU's realtime guarantees"?
I've tried hooking up prudebug, but I can't capture moments when it's glitching in there. I've got hundreds of megs of CSV files output from my logic analyzer, but all that's really told me is that every few hundred thousand bits, it glitches out for some reason. It sometimes remains stable for millions of bits. But then suddenly the PRU timing glitches out, but only ever by a few hundred nanoseconds. Enough to matter, but not so much as to make it obvious what's going wrong.