We take Darios trick and apply it to the limited input range.
As 80 can be divided by 16, we can, in order to round down to the next multiple of 80, discard the rightmost hex digit (set it to zero) and round down the number left of this digit to the next multiple of 5.
That can be done by determining the remainder of such a division and subtracting it. The leftmost digit is a constant 0xE. 0xE00 mod 5 = 4. The second digit is in the hexadecimal system multiplied by 16. 16 mod 5 = 1. So the remainder of the division by 5 is 4 + second digit + third digit.
As we have to shift the input bytes to get to the middle digits and shift back to subtract from the input (or as an alternate way, subtract to and from a shifted number and shift back the difference), we can also do our calculations with numbers shifted to the left, as long as they fit into a byte to save shift operations.
The maximum sum of the two middle digits and 4 is 4 + 0x6 + 0xf = 25. So we can calculate with numbers up to 8x as high to stay below 255.
There are different ways to get the remainder of a division by 5 from a number between 4 and 25. Either by a loop or by a series of tests of the range and branching. (Branches and memory accesses are cheap on those early processors compared to today's.) We have to find a compromise between execution time and code size.
Instead of tests in order to set the flags for branching, we can do actual calculations and branch dependent on the result, which also saves instructions.
The used flags are Carry, Zero and Sign.
Carry/Borrow gives us the information that the previous addition or subtraction went above 255 or below 0 and wrapped around.
Zero/Equal tells us that the result was 0.
Sign gives us the most significant bit, or that the previous sum is actually 16 or higher, if we do all the calculations multiplied by 8. 16*8=128, which is the value of the MSB of an 8 bit unsigned int.
Assuming the index register points to the high byte of the input number followed by the low byte in memory (big endian convention as was often used by Motorola, but the indices can be simply changed in the following code, when accessing memory).
LDAA #00H,X ; load high byte into A
ANDA #0FH ; take lower digit
LDAB #01H,X ; load low byte into B
ANDB #F0H ; select higher digit of B
ASLA ; do everything with numbers * 8
ASLA
ASLA
LSRB ; shift right by 1
ABA ; add B to A
ADDA #20H ; add 8*4 for contribution of 0xE000
AGAIN:
SUBA #28H ; subtract 8*5
BCC AGAIN ; no borrow, do it again
ADDA #28H ; we subtracted once too much, undo
ASLA ; multiply by 2 again
TAB ; transfer A to B
LDAA #01H,X ; load low byte into A
ANDA #F0H ; set lower digit to 0
SBA ; subtract B from A, keep carry
STAA #01H,X ; store low byte back
BCC FINISHED; no borrow occured
DEC #00H,X ; borrow -> decrement high byte
FINISHED:
This solution takes 34 bytes and executes up to 30 instructions (and minimally executes 20).
Variant 1:
LDAA #00H,X ; load high byte into A
ANDA #0FH ; take lower digit
LDAB #01H,X ; load low byte into B
ANDB #F0H ; select higher digit of B
ASLA ; do everything with numbers * 8
ASLA
ASLA
LSRB ; shift right by 1
ABA ; add B to A
BPL PLUS0_15; 0..15
SUBA #(21*8); 16..21 -21
BCC GOOD ; 21 change = -21
ADDA #(5*8) ; 16..20 -21+5
BRA GOOD ; change = -16
PLUS0_15: ; 0..15
BNE PLUS1_15; 1..15
ADDA #(4*8) ; 0 +4
BRA GOOD ; change = +4
PLUS1_15: ; 1..15
SUBA #(11*8); -11
BCC GOOD ; 11..15 change = -11
ADDA #(5*8) ; -11+5
BCS GOOD ; 6..10 change = -6
ADDA #(5*8) ; 1..5 -11+5+5
; change = -1
GOOD:
ASLA ; multiply by 2 again
TAB ; transfer A to B
LDAA #01H,X ; load low byte into A
ANDA #F0H ; set lower digit to 0
SBA ; subtract B from A, keep carry
STAA #01H,X ; store low byte back
BCC FINISHED; no borrow occured
DEC #00H,X ; borrow -> decrement high byte
FINISHED:
This solution takes 52 bytes and executes up to 24 instructions (and minimally executes 19). Faster, but larger.
Variant 2:
LDAA #00H,X ; load high byte into A
ANDA #0FH ; take lower digit
LDAB #01H,X ; load low byte into B
ANDB #F0H ; select higher digit of B
ASLA ; do everything with numbers * 8
ASLA
ASLA
LSRB ; shift right by 1
ABA ; add B to A
BPL PLUS0_15; 0..15
SUBA #(21*8); 16..21 -21
BRA SAMECODE
;BCC GOOD ; 21 change = -21
;ADDA #(5*8); 16..20 -21+5
;BRA GOOD ; change = -16
PLUS0_15: ; 0..15
CMPA #(6*8);
BCC PLUS6_15; 6..15
SUBA #(6*8) ; -1
BRA SAMECODE
;BCC GOOD ; 1..5 change = -1
;ADDA #(5*8); 0 -1+5
;BRA GOOD ; change = +4
PLUS6_15: ; 6..15
SUBA #(11*8); -11
SAMECODE:
BCC GOOD ; 11..15 change = -11
ADDA #(5*8) ; -11+5
GOOD:
ASLA ; multiply by 2 again
TAB ; transfer A to B
LDAA #01H,X ; load low byte into A
ANDA #F0H ; set lower digit to 0
SBA ; subtract B from A, keep carry
STAA #01H,X ; store low byte back
BCC FINISHED; no borrow occured
DEC #00H,X ; borrow -> decrement high byte
FINISHED:
This solution takes 46 bytes and executes up to 24 instructions (and minimally executes 20). A good bit smaller with code reuse, a bit worse optimal case, same worst case. One should better compare the average case.
Variant 3:
LDAA #00H,X ; load high byte into A
ANDA #0FH ; take lower digit
LDAB #01H,X ; load low byte into B
ANDB #F0H ; select higher digit of B
ASLA ; do everything with numbers * 8
ASLA
ASLA
LSRB ; shift right by 1
ABA ; add B to A
BPL PLUS0_15; 0..15
SUBA #(21*8); 16..21 -21
BCC GOODA ; 21 change = -21
BRA SAMECODE
;ADDA #(5*8); 16..20 -21+5
;BRA GOODA ; change = -16
PLUS0_15: ; 0..15
SUBA #(6*8) ;
BCS PLUS0_5 ; 0..5
TAB ; Transfer A to B (keep safe for 6..10)
SUBA #(5*8) ; -6-5
BCC GOODA ; 11..15 change = -11
BRA GOODB ; 6..10 change = -6
PLUS0_5: ; 0..5
ADDA #(5*8) ; -6+5
BCS GOODA ; 1..5 change = -1
SAMECODE:
ADDA #(5*8) ; 0 -6+5+5
; change = +4
GOODA:
TAB ; transfer A to B
GOODB:
ASLB ; multiply by 2 again
LDAA #01H,X ; load low byte into A
ANDA #F0H ; set lower digit to 0
SBA ; subtract B from A, keep carry
STAA #01H,X ; store low byte back
BCC FINISHED; no borrow occured
DEC #00H,X ; borrow -> decrement high byte
FINISHED:
This solution takes 51 bytes and executes up to 23 instructions (and minimally executes 19). Larger again, but even better worst case.
A more conventional solution (also working with other divisors than 0x50):
LDAA #00H,X ; load high byte
SUBA #DCH ; subtract 0xDC; 0xDC00 is divisible by 80; prevent overflow of counter, shorten execution time; we know input is at least 0xE000
CLR #00H,X ; clear counter
LDAB #01H,X ; load low byte
REP1:
INC #00H,X ; count
SUBB #50H ; try subtracting 0x50
SBCA #00H ; subract with borrow
BCC REP1 ; not finished
LDAA #DBH ; initialize high byte with 0xDB
LDAB #B0H ; initialize low byte with 0xB0 (counter is 1 too high)
REP2:
ADDB #50H ; add 0x50 to low byte
ADCA #00H ; add carry to high byte
DEC #00H,X ; decrease counter
BNE REP2 ; until zero
STAB #01H,X ; store back low byte
STAA #00H,X ; store back high byte
This solution needs 32 bytes and executes up to 312 instructions (minimum 112). At least smaller in size.
As comparison the approach with rounding down to multiples of 0x20 instead of of 0x50:
LDAA #01H,X ; load low byte
ANDA #E0H ; zero the 5 low bits
STAA #01H,X ; store back
would need 6 bytes and execute 3 instructions.