Author SHA1 Message Date
Michał Isalski ad8fd7e68e Finished bitblit function - to be tested! 2026-09-07 01:44:51 +02:00
ShatteredMINT 50f7040f59 Merge pull request 'Add mem.compare, mem.copy, mem.fill32, unsafe_mem.compare' (#13) from Mutex/symphony_stdlib:mem_ops into main
Reviewed-on: TCShenanigans/symphony_stdlib#13
2026-09-04 21:55:22 +02:00
ShatteredMINT 57934309c2 Merge branch 'main' into mem_ops 2026-09-04 21:53:34 +02:00
ShatteredMINT 32a157e87d Merge pull request 'Changed mul_low algorithm to radix-4 multiplication' (#16) from Micha_i/symphony_stdlib:fast-math into main
Reviewed-on: TCShenanigans/symphony_stdlib#16
2026-09-04 08:21:49 +02:00
Michał Isalski a00c7f2a84 Changed mul_low algorithm to radix-4 multiplication 2026-09-04 00:11:33 +02:00
MutexRaceCondition ae0684f98a Merged main into mem_ops 2026-09-03 21:15:08 +02:00
MutexRaceCondition 257735d776 Remove leftover files 2026-09-03 21:01:07 +02:00
MutexRaceCondition 809f6e35d5 Merged with main 2026-09-03 20:58:17 +02:00
ShatteredMINT c83269fcf7 fix path to stdlib.asm 2026-09-03 20:35:41 +02:00
ShatteredMINT 6bf0c0a4fc add examples directory 2026-09-03 20:34:56 +02:00
ShatteredMINT 697445feec add tests folder 2026-09-03 20:34:16 +02:00
ShatteredMINT 7c49c5212e move assembly files into src directory 2026-09-03 20:32:40 +02:00
ShatteredMINT 6ff1a43c32 Merge pull request 'Added find_index array function' (#9) from Micha_i/symphony_stdlib:array-functions into main
Reviewed-on: TCShenanigans/symphony_stdlib#9
2026-09-03 20:30:09 +02:00
Michał Isalski b5249a43b9 Merge branch 'main' into array-functions 2026-09-03 20:29:24 +02:00
MutexRaceCondition 655ce0237e Fixed unsafe_mem.compare not working correctly when both input pointers are the same 2026-09-03 18:48:57 +02:00
MutexRaceCondition a1e19112c5 Specialised unsafe.asm to unsafe_mem.asm, and renamed unsafe.memcmp to unsafe.compare 2026-09-03 15:22:40 +02:00
MutexRaceCondition 3d874cebbd Merged with main 2026-09-03 15:17:00 +02:00
MutexRaceCondition 4aca17aa40 Moved memory operations into mem.asm (and renamed them) 2026-09-03 15:15:54 +02:00
ShatteredMINT 0188969dab Merge pull request 'Added read_line console function' (#10) from Micha_i/symphony_stdlib:console-functions into main
Reviewed-on: TCShenanigans/symphony_stdlib#10
2026-09-03 14:56:12 +02:00
Micha_i 5c9a33c24d Merge branch 'main' into array-functions 2026-09-03 09:04:11 +02:00
Micha_i e4edb56d08 Merge branch 'main' into console-functions 2026-09-03 09:04:03 +02:00
Michał Isalski aa8eef7b1f Fixed at-address 2026-09-02 22:56:24 +02:00
ShatteredMINT 37d0e556d2 Merge pull request 'clean up confusion about function template' (#14) from meta-documentation into main
Reviewed-on: TCShenanigans/symphony_stdlib#14
2026-09-02 20:49:38 +02:00
ShatteredMINT 26d3a4d4f4 clean up confusion about function template 2026-09-02 20:49:22 +02:00
Michał Isalski 870b6ebc39 Merge branch 'main' into console-functions 2026-09-02 20:41:38 +02:00
Michał Isalski 1e4ad8f64c Merge branch 'main' into array-functions 2026-09-02 20:40:49 +02:00
MutexRaceCondition 7b43e75a9d Added memcmp, memcpy, memset32, unsafe.memcmp 2026-09-02 19:41:56 +02:00
Michal Isalski ca1ba67fd8 Added a shift conversion LUT and finished read_line (except Writeback) 2026-09-02 13:20:55 +02:00
ShatteredMINT 786adfdb1c Merge pull request 'Changed docs of math functions to conform to new guidelines' (#12) from Micha_i/symphony_stdlib:math-functions-docs into main
Reviewed-on: TCShenanigans/symphony_stdlib#12
2026-09-02 11:46:14 +02:00
Michał Isalski f9bb69f1f5 Added storing shift status 2026-09-02 02:03:00 +02:00
Michal Isalski 20e1f17181 Fixed swapped pop instructions 2026-09-01 13:52:45 +02:00
Michal Isalski 625f01166a Made find_index conform to new doc guidelines 2026-09-01 13:52:26 +02:00
Michal Isalski 65caf8aff9 Made read_line compliant to new doc guidelines 2026-09-01 13:47:09 +02:00
Michal Isalski 1dca814388 One instruction less by using another register 2026-09-01 10:54:29 +02:00
Michal Isalski 9dfb8248a6 Added handling for skipping non-renderable characters
Added handling for backspace character
2026-09-01 10:51:31 +02:00
Michał Isalski 4b962a1f87 Added read_line function 2026-09-01 00:38:37 +02:00
Michał Isalski e6a0739e95 pleegwat's code review fixes 2026-08-31 23:57:55 +02:00
Michał Isalski 1885f304ad Merge branch 'array-functions' of https://gitea.shatteredmint.net/Micha_i/symphony_stdlib into array-functions 2026-08-31 23:55:16 +02:00
Michał Isalski 63b37ab094 Added comment about predicate context 2026-08-31 23:06:59 +02:00
Michał Isalski 6981ff027d Tested and fixed stride->shift conversion missing 2026-08-31 23:06:59 +02:00
Michal Isalski 88c91a8eec Reduced operations to get -1 in register 2026-08-31 23:06:59 +02:00
Michal Isalski a02f0646b4 Fixed the predicate return address 2026-08-31 23:06:59 +02:00
Michal Isalski 7abfcabd73 Added find_index array function 2026-08-31 23:06:59 +02:00
Michał Isalski b476b8aaa3 Added comment about predicate context 2026-08-31 23:06:40 +02:00
Michał Isalski bddb153d9b Tested and fixed stride->shift conversion missing 2026-08-31 23:03:23 +02:00
Michal Isalski 29103d4335 Reduced operations to get -1 in register 2026-08-31 15:30:11 +02:00
Michal Isalski 837e8ba0a6 Fixed the predicate return address 2026-08-31 13:52:27 +02:00
Michal Isalski 6ca77970b8 Added find_index array function 2026-08-31 13:44:07 +02:00
14 changed files with 792 additions and 18 deletions
+2 -2
View File
@@ -16,8 +16,8 @@ All functions in the standard library should follow the following outline:
; Arguments: <which register contains what argument>
; Result: <what is the result, and where is it stored>
; Clobbers: <list of registers that are clobbered>
fn_label: <;SHOULD BE INLINED>
CODE
<label>: <;SHOULD BE INLINED>
<CODE>
```
+1 -1
View File
@@ -3,7 +3,7 @@
This is a standard library for symphony.
It is both intended as a practical toolkit to develop more complex software as well as a teaching resource.
If you just want to use the standard library [[stdlib.asm]] is your main header, include it after your code.
If you just want to use the standard library [[src/stdlib.asm]] is your main header, include it after your code.
If you are using it as a learning resource have a look at the [teaching folder](teaching).
+3
View File
@@ -0,0 +1,3 @@
# Examples
Examples of how to use the standard library to accomplish a task.
+100
View File
@@ -0,0 +1,100 @@
@0x10000 ; Example address until we get a proper memory map for this
; Shift conversion table
; It stores the mapping from value 32-127 of the ASCII table to their shifted equivalents (both ways) in the standard US keyboard layout
; e.g. 1 -> !
U8 32 ; Space -> Space
U8 49 ; ! -> 1
U8 39 ; " -> '
U8 51 ; # -> 3
U8 52 ; $ -> 4
U8 53 ; % -> 5
U8 55 ; & -> 7
U8 34 ; ' -> "
U8 57 ; ( -> 9
U8 48 ; ) -> 0
U8 56 ; * -> 8
U8 61 ; + -> =
U8 60 ; , -> <
U8 95 ; - -> _
U8 62 ; . -> >
U8 63 ; / -> ?
U8 41 ; 0 -> )
U8 33 ; 1 -> !
U8 64 ; 2 -> @
U8 35 ; 3 -> #
U8 36 ; 4 -> $
U8 37 ; 5 -> %
U8 94 ; 6 -> ^
U8 38 ; 7 -> &
U8 42 ; 8 -> *
U8 40 ; 9 -> (
U8 59 ; : -> ;
U8 58 ; ; -> :
U8 44 ; < -> ,
U8 43 ; = -> +
U8 46 ; > -> .
U8 47 ; ? -> /
U8 50 ; @ -> 2
U8 97 ; A -> a
U8 98 ; B -> b
U8 99 ; C -> c
U8 100; D -> d
U8 101; E -> e
U8 102; F -> f
U8 103; G -> g
U8 104; H -> h
U8 105; I -> i
U8 106; J -> j
U8 107; K -> k
U8 108; L -> l
U8 109; M -> m
U8 110; N -> n
U8 111; O -> o
U8 112; P -> p
U8 113; Q -> q
U8 114; R -> r
U8 115; S -> s
U8 116; T -> t
U8 117; U -> u
U8 118; V -> v
U8 119; W -> w
U8 120; X -> x
U8 121; Y -> y
U8 122; Z -> z
U8 123; [ -> {
U8 124; \ -> |
U8 125; ] -> }
U8 125; ^ -> 6
U8 45 ; _ -> -
U8 126; ` -> ~
U8 65 ; a -> A
U8 66 ; b -> B
U8 67 ; c -> C
U8 68 ; d -> D
U8 69 ; e -> E
U8 70 ; f -> F
U8 71 ; g -> G
U8 72 ; h -> H
U8 73 ; i -> I
U8 74 ; j -> J
U8 75 ; k -> K
U8 76 ; l -> L
U8 77 ; m -> M
U8 78 ; n -> N
U8 79 ; o -> O
U8 80 ; p -> P
U8 81 ; q -> Q
U8 82 ; r -> R
U8 83 ; s -> S
U8 84 ; t -> T
U8 85 ; u -> U
U8 86 ; v -> V
U8 87 ; w -> W
U8 88 ; x -> X
U8 89 ; y -> Y
U8 90 ; z -> Z
U8 91 ; { -> [
U8 92 ; | -> \
U8 93 ; } -> ]
U8 96 ; ~ -> `
+77
View File
@@ -0,0 +1,77 @@
; Returns the index of the first element matching the provided predicate function (or -1 if not found)
; Arguments:
; r1 - The array pointer
; r2 - The array length (number of items)
; r3 - The stride (size of one item) - either 1, 2 or 4 (bytes)
; r4 - The predicate
; r5 - Predicate context
; Result:
; r1 - The index of the first element matching the provided predicate function (or -1 if not found)
; Clobbers: r2, r3, r4, r5, r6, + what the predicate clobbers
; Info:
; The predicate function should follow the stdlib calling convention
; The predicate receives two arguments (the value and the predicate context) and should return either a zero when the value is not the one we search for
; , or any other result if it is the searched-for item.
pub find_index:
push r12 ; We will store the predicate pointer here
push r11 ; We will store the current pointer here
push r10 ; We will store the stride here
push r9 ; We will store the final address here
push r8 ; We will store the mask here
mov r12, r4
mov r11, r1
mov r10, r3
mov r9, r2
lsr r6, r3, 1 ; We turn the stride into a byte shift
lsl r9, r9, r6 ; We calculate bytes left
add r9, r9, r1 ; We add the start address to get the final address
push r13 ; We save up the return address because we will provide our own to the predicate
push r1 ; We need the array pointer to calculate the item index
counter r13
add r13, r13, 52 ; Point to just after the predicate call - we can set this up now so we don't waste loop cycles
nand r8, zr, zr ; We create a mask of 0xFFFFFFFF
mov r6, 4
sub r6, r6, r3 ; We create a "negative stride", e.g. 4 -> 0, 2 -> 2, 1 -> 3
lsl r6, r6, 3
lsr r8, r8, r6 ; We shift the mask by the negative stride to obtain the proper mask for a value
; e.g. stride 4 -> mask is 0xFFFFFFFF
; stride 2 -> mask is 0x0000FFFF
; stride 1 -> mask is 0x000000FF
push r5 ; We save the predicate context on the stack
find_index_loop:
load_32 r1, [r11] ; We load the element
and r1, r1, r8 ; We mask it to handle stride 2 and 1 cases
load_32 r2, [sp] ; We load the predicate context into r2
jmp r12 ; We call the predicate
cmp r1, zr
jne find_index_found_item ; If we found the item, we jump out
; If we didn't, move to next item
add r11, r11, r10 ; We add the stride to the pointer
cmp r11, r9 ; We compare with the final address
jne find_index_loop ; If we did not reach the end we jump back into the loop
find_index_not_found:
add sp, sp, 8 ; The predicate context and old array pointer are not useful
pop r13 ; We get our return address
nand r1, zr, zr ; We put -1 in r1
jmp find_index_postamble
find_index_found_item:
add sp, sp, 4 ; The predicate context is not useful
pop r1 ; We get the array pointer
pop r13 ; We get our return address
sub r1, r11, r1 ; We calculate the bytes from the start
lsr r10, r10, 1 ; We shift the stride to get the amount to shift the bytes for
lsr r1, r1, r10 ; We shift to get the index of the item
find_index_postamble:
pop r8
pop r9
pop r10
pop r11
pop r12
jmp r13 ; Return
View File
+69
View File
@@ -0,0 +1,69 @@
; Reads a line from the keyboard and fills the specified buffer with it
; Does not support Shift or any other special keys
; Arguments:
; r1 - pointer to the buffer
; Result:
; r1 - pointer to the same buffer
; Clobbers: r2, r3, r4, r5, r6, r7
pub read_line:
mov r6, 1
lsl r6, r6, 16
sub r6, r6, 32 ; Calculating the address to the shift LUT
mov r2, 0 ; Storing the shift status here
mov r4, 0 ; Storing the last key here, so we don't repeat the same key
mov r3, r1 ; The pointer to after the last character
read_line_keyloop:
keyboard r5
cmp r5, r4
je read_line_keyloop ; If the current key is same as previous, we loop
mov r4, r5 ; Storing current key as previous
cmp r5, 0x120 ; Is the key renderable or special?
jb read_line_special ; If the key was special, we handle it separately
xor r5, r5, 0x100 ; Clearing the "down" bit
cmp r2, zr ; Checking for shift status
je read_line_store ; If shift is up, we skip conversion
;;; converting from shift-down to shift-up keys
add r7, r6, r5 ; Calculating the index of the shift conversion
load_8 r5, [r7] ; Loading the shifted value
;;;
read_line_store:
store_8 [r3], r5 ; Else, we store the key in the buffer
add r3, r3, 1 ; We advance forward
; TODO: Writeback
jmp read_line_keyloop
read_line_special:
cmp r5, 13 ; Was the key Backspace?
je read_line_backspace ; If yes we need to move one character back
cmp r5, 10 ; Was the key Enter?
je read_line_finished ; If so, we're finished
and r5, r5, 0x1FB ; Mask out the left/right shift direction bit
cmp r5, 0x110 ; Was the key Shift Down?
and flags, flags, 0x1 ; We care only about equality bit
or r2, r2, flags ; If shift was down before, it still is. If it was pressed now, it is down now
cmp r5, 0x010 ; Was the key Shift Up?
and flags, flags, 0x1 ; We care only about equality bit
xor flags, flags, 0x1 ; We invert it, i.e. "if it's not up"
and r2, r2, flags ; The shift can be kept down if it's not currently up
jmp read_line_keyloop ; If no special handling, we loop back
read_line_backspace:
cmp r3, r1 ; Compare the current pointer to start of buffer
je read_line_keyloop ; If we are at the start, we loop
sub r3, r3, 1 ; We move back one character
; TODO: Writeback
jmp read_line_keyloop
read_line_finished:
store_8 [r3], zr ; We store null at the end so the string is finished
jmp r13
+256
View File
@@ -0,0 +1,256 @@
include errno
; Performs a Bitblock transfer, copying a section of the Source to the Destination, applying a given Mode operation
; Arguments:
; r1 - Destination Buffer pointer
; r2 - Source Buffer pointer
; r3 - X destination coordinate
; r4 - Y destination coordinate
; r5 - X source coordinate
; r6 - Y source coordinate
; r7 - Mode:
; 000 - SRCCOPY (copies source over destination)
; 001 - XOR
; 010 - AND
; 011 - NOT
; 100 - OR
; 101 - reserved
; 110 - reserved
; 111 - SRCALPHA (copies source over destination if: source != 0 (for 8bpp), source.alpha != 0 (for 32bpp))
; Stack argument 1: Width (in bytes)
; Stack argument 2: Height (in bytes)
; Result: None
; Clobbers: r1, r2, r3, r4, r5, r6, r7
; Errors: BUFFER_DEPTH_MISMATCH - if the buffers have different bit depths
pub bitblit:
push r8
push r9
; Checking if the bit depths are correct
add r8, r1, 6 ; r8 = pointer to destination depth
load_16 r8, [r8] ; r8 = dest depth
add r9, r2, 6 ; r9 = pointer to source depth
load_16 r9, [r9] ; r9 = src depth
cmp r8, r9
je bitblit_depth_good
; Bit depths differ - error
pop r9
pop r8
add sp, sp, 8 ; Removing the stack Arguments
mov flags, errno.BUFFER_DEPTH_MISMATCH
jmp r13
bitblit_depth_good:
push r8
add sp, sp, 12
load_32 r8, [sp] ; r8 = height
add sp, sp, 4
load_32 r9, [sp] ; r9 = width
sub sp, sp, 16
push r10
push r11
push r12
push r13
add r10, r1, 4 ; r10 = pointer to destination stride
load_16 r10, [r10] ; r10 = destination stride
add r11, r2, 4 ; r11 = pointer to source stride
load_16 r11, [r11] ; r11 = source stride
; Calculating destination index
push r10
mov r12, 0
bitblit_destination_index:
cmp r4, zr
je bitblit_after_destination_index
mov flags, r4
jne bitblit_destination_index_afteradd
add r12, r12, r10
bitblit_destination_index_afteradd:
lsr r4, r4, 1
lsl r10, r10, 1
jmp bitblit_destination_index
bitblit_after_destination_index:
add r4, r4, r12 ; Now r4 is dest_y * dest_stride
load_32 r1, [r1] ; r1 = destination data pointer
add r1, r1, r4 ; Now r1 is dest_pointer + dest_y * dest_stride
add r1, r1, r3 ; Now r1 is dest_pointer + dest_y * dest_stride + dest_x
pop r10
push r11
; Calculating source index
mov r12, 0
bitblit_source_index:
cmp r6, zr
je bitblit_after_source_index
mov flags, r6
jne bitblit_source_index_afteradd
add r12, r12, r11
bitblit_source_index_afteradd:
lsr r6, r6, 1
lsl r11, r11, 1
jmp bitblit_source_index
bitblit_after_source_index:
add r6, r6, r12 ; Now r6 is src_y * src_stride
load_32 r2, [r2] ; r2 = source data pointer
add r2, r2, r6 ; Now r2 is src_pointer + src_y * src_stride
add r2, r2, r5 ; Now r2 is src_pointer + src_y * src_stride + src_x
pop r11
; r3, r4, r5, r6 are now free
mov r3, bitblit_mode
mov r12, bitblit_end_mode
lsl r7, r7, 3 ; 2 instructions per mode
add r12, r7, r12
add r7, r7, r3
; r1 is line dest pointer (or bit depth)
; r2 is line src pointer (or temp)
; r3 is width iterator
; r4 is current dest pointer
; r5 is current src pointer
; r6 is current src value
; r7 is 32bit mode jump table destination
; r8 is height iterator
; r9 is width of bitblit
; r10 is dest stride
; r11 is src stride
; r12 is 8bit mode jump table destination
; r13 is current dest value
bitblit_copy_height:
cmp r8, zr
je bitblit_finish
bitblit_copy_line:
mov r3, 0
mov r4, r1 ; r4 = dest
mov r5, r2 ; r5 = src
push r1
add r1, sp, 20 ; Pointing to bit depth saved on stack
load_32 r1, [r1] ; r1 = bit depth
push r2
bitblit_copy_line_loop: ; Main loop that copies 4 bytes at a time
add r3, r3, 4
cmp r9, r3
je bitblit_copy_line_finish
jb bitblit_copy_line_end
load_32 r6, [r5] ; r6 = [src]
load_32 r13, [r4] ; r13 = [dest]
jmp r7
bitblit_mode:
nop ; SRCCOPY
jmp bitblit_save_bits
xor r6, r6, r13 ; XOR
jmp bitblit_save_bits
and r6, r6, r13 ; AND
jmp bitblit_save_bits
not r6, r6 ; NOT
jmp bitblit_save_bits
or r6, r6, r13 ; OR
jmp bitblit_save_bits
nop ; reserved
jmp bitblit_save_bits
nop ; reserved
jmp bitblit_save_bits
cmp r1, 8 ; SRCALPHA
jne bitblit_srcalpha_32 ; If bit depth is 32, we jump to alpha checking
; If bit depth is 8, we need to repack this value
mov r2, 0 ; our mask
lsr flags, r6, 24
cmp flags, zr
jne bitblit_srcalpha_8_1
or r2, r2, 0xFF
bitblit_srcalpha_8_1:
lsl r2, r2, 8
lsr flags, r6, 16
and flags, flags, 0xFF
cmp flags, zr
jne bitblit_srcalpha_8_2
or r2, r2, 0xFF
bitblit_srcalpha_8_2:
lsl r2, r2, 8
lsr flags, r6, 8
and flags, flags, 0xFF
cmp flags, zr
jne bitblit_srcalpha_8_3
or r2, r2, 0xFF
bitblit_srcalpha_8_3:
lsl r2, r2, 8
and flags, r6, 0xFF
cmp flags, zr
jne bitblit_srcalpha_8_4
or r2, r2, 0xFF
bitblit_srcalpha_8_4:
and r13, r13, r2
not r2, r2
and r6, r6, r2
or r6, r6, r13
jmp bitblit_save_bits
bitblit_srcalpha_32:
and flags, r6, 0xFF
cmp flags, zr
jne bitblit_save_bits ; If alpha channel is not zero, we save the value
mov r6, r13 ; If it was zero, we take [dest], i.e. don't overwrite
bitblit_save_bits:
store_32 [r4], r6
add r4, r4, 4
add r5, r5, 4
jmp bitblit_copy_line_loop
bitblit_copy_line_end: ; Now we need to handle up to 3 bytes that were not copied
sub r3, r3, 4
bitblit_copy_line_end_loop:
load_8 r6, [r5] ; r6 = [src]
load_8 r13, [r4] ; r13 = [dest]
jmp r12
bitblit_end_mode:
nop ; SRCCOPY
jmp bitblit_save_bits_end
xor r6, r6, r13 ; XOR
jmp bitblit_save_bits_end
and r6, r6, r13 ; AND
jmp bitblit_save_bits_end
not r6, r6 ; NOT
jmp bitblit_save_bits_end
or r6, r6, r13 ; OR
jmp bitblit_save_bits_end
nop ; reserved
jmp bitblit_save_bits_end
nop ; reserved
jmp bitblit_save_bits_end
cmp r6, zr ; SRCALPHA (will happen only in 8bpp)
jne bitblit_save_bits_end
mov r6, r13 ; [src] = 0, so we take [dest], i.e. don't overwrite
bitblit_save_bits_end:
store_8 [r4], r6
add r3, r3, 1
add r4, r4, 1
add r5, r5, 1
cmp r9, r3
ja bitblit_copy_line_end_loop
bitblit_copy_line_finish:
pop r2
pop r1
add r1, r1, r10
add r2, r2, r11 ; Moving dest & src to next line
sub r8, r8, 1
jmp bitblit_copy_height
bitblit_finish:
pop r13
pop r12
pop r11
pop r10
add sp, sp, 4 ; Bit depth skip
pop r9
pop r8
add sp, sp, 8
mov flags, errno.OK
jmp r13
+20 -11
View File
@@ -6,20 +6,29 @@
; r1 - The lower 32 bits of the result
; Clobbers: r2, r3, r4, r5
pub mul_low:
mov r3, 0 ; result
mov r4, 31 ; loop counter
add r3, r2, r2 ; r3 = r2 * 2
mov r4, 0 ; r4 has the result
mul_low_loop:
asr r5, r2, 31
and r5, r5, r1
lsl r5, r5, r4
add r3, r3, r5
lsl r2, r2, 1
sub r4, r4, 1
cmp r4, 0
jge mul_low_loop
mov r1, r3
and r5, r1, 1 ; a0
neg r5, r5 ; mask
and r5, r2, r5 ; a0 ? r2 : 0
add r4, r4, r5
and r5, r1, 2 ; a1 is now 0 or 2
lsr r5, r5, 1 ; normalize to 0/1
neg r5, r5 ; mask
and r5, r3, r5 ; a1 ? r2 * 2 : 0
add r4, r4, r5
lsl r2, r2, 2
lsl r3, r3, 2
lsr r1, r1, 2
cmp r1, zr
jne mul_low_loop
mov r1, r4
jmp r13
; Calculates the absolute value of the value provided in the r1 register
+174
View File
@@ -0,0 +1,174 @@
; int compare(uint8_t* a, uint8_t* b, size_t count);
; Compares two memory segment of equal length lexicographically.
;
; Arguments:
; - `r1`: A pointer to the first memory segment.
; - `r2`: A pointer to the second memory segment.
; - `r3`: The size of both memory segments.
; Results:
; - `r1`:
; - `0` if both segments are equal.
; - `<0` if the first segment is less than the second segment.
; - `>0` if the first segment is greater than the second segment.
;
pub compare:
; Exclusive end point of the first segment.
add r3, r3, r1
sub r3, r3, 4
_compare__loop:
load_32 r4, [r1]
add r1, r1, 4
load_32 r5, [r2]
add r2, r2, 4
; Comparing two sequences of 4 bytes lexicographically is equivalent to
; comparing the corresponding big endian 32 bit words.
cmp r4, r5
jne _compare__break
; Check if there are enough bytes left to continue with the vectorized loop.
cmp r1, r3
jbe _compare__loop
; `r3 + 4 - r1 = <remaining byte count> = r3 - r1 mod 4`
sub flags, r3, r1
; Check if one of the lowest 2 bits is non-zero
jbe _compare__rem
; If not, we are done. Both segments are equal.
mov r1, 0
jmp r13
_compare__break:
; `flags` is the comparison result in the format of `cmp`. Convert it to the desired format.
; 00 => 0x40000000 > 0
; 01 => 0x00000000 = 0
; 10 => 0xC0000000 < 0
xor r1, flags, 1
lsl r1, r1, 30
jmp r13
_compare__rem:
; Compute `S = 8*(4 - <remaining byte count>)` and
; [r1] >> S, [r2] >> S
mov r3, 8
load_32 r4, [r1]
sub r3, r3, flags
load_32 r5, [r2]
lsl r3, r3, 3
lsr r4, r4, r3
lsr r5, r5, r3
; Compare both values, now with garbage bytes removed.
cmp r4, r5
jmp _compare__break
; void copy(void* src, void* dest, size_t count);
; Copies `count` bytes from `src` to `dest`. The two memory segments must not overlap.
;
; Arguments:
; - `r1`: Pointer to the memory segment to be copied.
; - `r2`: Pointer to the memory segment to be copied into.
; - `r3`: Byte size of both the `src` and `dest` segments.
;
pub copy:
; Exclusive end point of the source segment.
add r3, r1, r3
; Last index from where we can safely copy 8 bytes per loop iteration.
sub r3, r3, 8
jmp _copy__loop_entry
_copy__loop:
; Copy 8 bytes from `src` to `dest`.
load_32 flags, [r1]
add r1, r1, 4
store_32 [r2], flags
add r2, r2, 4
load_32 flags, [r1]
add r1, r1, 4
store_32 [r2], flags
add r2, r2, 4
_copy__loop_entry:
; Check if we can process more data in the vectorized loop.
cmp r1, r3
jbe _copy__loop
; The remaining amount of bytes `R` is `R = r3 + 8 - r1 = r3 - r1 mod 8`.
sub flags, r3, r1
; Test if `R` is not a multiple of `4`, i.e. the lowest 2 bits are non-zero.
jbe _copy__rem
; `R` is a multiple of `4`. Special case this.
; Check if `R` is `0`, i.e. the third bit is also 0. In that case, we are already done.
; There are no conditional indirect jumps, so we can't return immediately.
jge _copy__ret
; `R = 4`. No need to update `r1` or `r2`, we don't need them anymore.
load_32 flags, [r1]
store_32 [r2], flags
_copy__ret:
; Return
jmp r13
_copy__rem:
; Optimize the remaining cases for code size.
; End point of the source segment.
add r3, r3, 8
; We already handled the case `R = 0` earlier,
; so no bounds check needed for the first iteration.
_copy__rem_loop:
; Copy 1 byte.
load_8 flags, [r1]
add r1, r1, 1
store_8 [r2], flags
add r2, r2, 1
; Check if we are still within the bounds.
cmp r1, r3
jb _copy__rem_loop
jmp r13
; void fill32(uint8_t* dest, size_t count, uint32_t value);
; Fills `count` bytes in `dest` with `value`. If `count` is not a multiple of 4,
; the least significant bytes of `value` are cut off for the last entry.
;
; Arguments:
; - `r1`: A pointer to the destination segment.
; - `r2`: The size of the destination segment.
; - `r3`: The 32 bit value that the segment is filled with.
;
pub fill32:
; Exclusive end point of the destination segment.
add r2, r1, r2
; Last index from where we can safely write 8 bytes per loop iteration.
sub r2, r2, 8
jmp _fill32__entry
_fill32__loop:
; Set 8 bytes per loop iteraion.
store_32 [r1], r3
add r1, r1, 4
store_32 [r1], r3
add r1, r1, 4
_fill32__entry:
; Check if we can process more data in the vectorized loop.
cmp r1, r2
jbe _fill32__loop
; The remaining amount of bytes `R` is `R = r2 + 8 - r1 = r2 - r1 mod 8`.
sub flags, r2, r1
; Check if the third bit of the remainder is cleared.
jge _fill32__r4
; Otherwise set 4 bytes.
store_32 [r1], r3
add r1, r1, 4
_fill32__r4:
; Check if the two least significant bits of the remainder are zero.
ja _fill32__ret
; Handle the remaining bytes `R` individually, in reverse order.
add r2, r2, 4
; `r1 + 4 - r2 = 4 - R`.
sub flags, r1, r2
; Exclusive end point of the destination segment.
add r2, r2, 4
; Shift out the least significant `8*(4 - R)` bits of the value.
lsl flags, flags, 3
lsr r3, r3, flags
jmp _fill32__loop2_entry
_fill32__loop2:
sub r2, r2, 1
; Write the least significant byte of the value...
store_8 [r2], r3
; and then shift it out.
lsr r3, r3, 8
_fill32__loop2_entry:
cmp r1, r2
jb _fill32__loop2
_fill32__ret:
jmp r13
+8
View File
@@ -0,0 +1,8 @@
pub include bit
pub include imath
pub include array
pub include console
pub include mem
; Needs to be last!
pub include LUTs
+77
View File
@@ -0,0 +1,77 @@
; int compare(uint8_t* a, uint8_t* b, size_t count);
; Compares two memory segment of equal length lexicographically.
; Temporarily modifies the byte at address `a + count`.
;
; Arguments:
; - `r1`: A pointer to the first memory segment.
; - `r2`: A pointer to the second memory segment.
; - `r3`: The size of both memory segments.
; Results:
; - `r1`:
; - `0` if both segments are equal.
; - `<0` if the first segment is less than the second segment.
; - `>0` if the first segment is greater than the second segment.
pub compare:
cmp r1, r2
je _compare__is_eq
; Exclusive end point of the second segment.
add r4, r3, r2
; Exclusive end point of the first segment.
add r3, r3, r1
load_8 r6, [r3]
load_8 flags, [r4]
; Check if the first bytes behind the sequences are equal.
cmp flags, r6
jne _compare__loop
; Change the byte directly behind the first segment.
xor r4, r6, 1
; This would be problematic if someone calls compare with a first segment
; whose end point overlaps the program memory of this function.
store_8 [r3], r4
_compare__loop:
load_32 r4, [r1]
add r1, r1, 4
load_32 r5, [r2]
add r2, r2, 4
; Comparing two sequences of 4 bytes lexicographically is equivalent to
; comparing the corresponding big endian 32 bit words.
cmp r4, r5
je _compare__loop
; We overshot in the loop; decrement r1 again. (Only by 2, we backtrack the rest if necessary later)
sub r1, r1, 2
; Restore the byte we changed.
store_8 [r3], r6
; We encountered two different words. Figure out what byte they differ on.
xor r4, r4, r5
; Store the flags for later, to figure out the return value.
mov r5, flags
; Check if at least one of the two most significant bytes is not 0.
cmp r4, 0xffff
jbe _compare__low2
; If it is, backtrack the remaining 2 indices.
; Shift the most significant bytes to the least significant ones.
sub r1, r1, 2
lsr r4, r4, 16
_compare__low2:
; r1 now points to a non-zero 16 bit value.
; If the 16 bit value at r1-2 is in-bounds, then it is 0.
; Check if the most significant byte of the 16 bit value is 0.
cmp r4, 0xff
ja _compare__low1
; If it is, our target is the least significant byte.
add r1, r1, 1
_compare__low1:
; Otherwise, the target is that non-zero byte.
; Check if the target is out of bounds, i.e. the loop terminated through the "bounds check".
cmp r1, r3
jae _compare__is_eq
; r5 is the comparison result in the format of `cmp`. Convert it to the desired format.
; 00 => 0x40000000 > 0
; 01 => 0x00000000 = 0
; 10 => 0xC0000000 < 0
xor r1, r5, 1
lsl r1, r1, 30
jmp r13
_compare__is_eq:
mov r1, 0
jmp r13
-2
View File
@@ -1,2 +0,0 @@
pub include bit
pub include imath
+3
View File
@@ -0,0 +1,3 @@
# Tests
Tests for the standard library go here, tests are allowed to depend on the recommended spec.isa changes.