Skip to content

Lab 29: Your First Assembly Function

Time: ~50 minutes | Prerequisites: Lab 28 | Hardware: Pico 2

Time to meet the machine

Echo waving welcome Assembly is much less frightening than it sounds. About a dozen instructions cover most of it, and you already know what all of them do. Time to transform!

What You'll Build

Assembly functions you can call from Python: adding two numbers, a counting loop, and summing an array 55x faster.

Learning Objectives

  • Write a function with @micropython.asm_thumb
  • Use registers r0-r7 as your only variables
  • Build a loop from a label, a compare and a branch
  • Read memory with ldr and address arithmetic
  • Measure the speedup over equivalent Python

Concepts Introduced

ID Concept
490 Inline Assembler
491 asm_thumb Decorator
492 CPU Register
493 General Purpose Register
494 Register Allocation
495 Move Instruction
496 Add Instruction
497 Compare Instruction
498 Conditional Branch
499 Assembly Label
500 Assembly Loop
501 Argument Passing Convention
502 Return Value Register
503 Machine Code
504 Instruction Mnemonic

Procedure

Open 29-first-assembly.py and work through it section by section:

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
# Lab 29: Your First Assembly Function
#
# Assembly is where you stop describing WHAT you want and start naming the
# exact instructions the CPU executes. There is no interpreter, no object
# layer, no type checking. Just registers and operations on them.
#
# It is much less frightening than it sounds. There are about a dozen
# instructions you need, and you already know what all of them do.

import machine
import micropython

machine.mem32[0xE000EDFC] = machine.mem32[0xE000EDFC] | (1 << 24)
machine.mem32[0xE0001000] = machine.mem32[0xE0001000] | 1


def rd():
    return machine.mem32[0xE0001004]


# =========================================================================
# PART 1 -- the smallest possible assembly function
# =========================================================================
print("=== PART 1: add two numbers ===")


@micropython.asm_thumb
def add_two(r0, r1):
    # Arguments arrive in r0, r1, r2, r3 -- in that order.
    # Whatever is in r0 when the function ends is the return value.
    add(r0, r0, r1)


print("add_two(40, 2) =", add_two(40, 2))
print("add_two(-5, 8) =", add_two(-5, 8))
print()
print("Two rules cover almost everything:")
print("  * arguments arrive in r0, r1, r2, r3")
print("  * whatever is in r0 at the end is returned")

# =========================================================================
# PART 2 -- registers are the only variables
# =========================================================================
print()
print("=== PART 2: registers ===")
print()
print("A Cortex-M33 has 16 core registers. You may freely use r0-r7.")
print("There is no 'x = 5' -- there is 'put 5 into register 3'.")
print()


@micropython.asm_thumb
def arithmetic_demo(r0):
    mov(r1, 10)          # r1 = 10
    mov(r2, 3)           # r2 = 3
    add(r3, r1, r2)      # r3 = 13
    sub(r4, r1, r2)      # r4 = 7
    mul(r1, r2)          # r1 = r1 * r2 = 30
    add(r0, r3, r4)      # r0 = 20
    add(r0, r0, r1)      # r0 = 50


print("arithmetic_demo() =", arithmetic_demo(0), " (expect 50)")

# =========================================================================
# PART 3 -- loops are compare-and-branch
# =========================================================================
print()
print("=== PART 3: loops ===")
print()
print("There is no 'for' or 'while'. A loop is a label, some work, a")
print("comparison, and a conditional jump back.")
print()


@micropython.asm_thumb
def sum_to_n(r0):
    # Add up 1 + 2 + ... + n
    mov(r1, 0)                  # total
    mov(r2, 1)                  # counter

    label(LOOP)
    add(r1, r1, r2)             # total += counter
    add(r2, 1)                  # counter += 1
    cmp(r2, r0)                 # compare counter with n
    ble(LOOP)                   # branch back while counter <= n

    mov(r0, r1)                 # return total


for n in (5, 10, 100):
    expected = n * (n + 1) // 2
    got = sum_to_n(n)
    print("sum_to_n(%3d) = %6d   (expect %6d)  %s"
          % (n, got, expected, "ok" if got == expected else "WRONG"))

# =========================================================================
# PART 4 -- reading memory
# =========================================================================
print()
print("=== PART 4: working with arrays ===")
print()
print("Assembly cannot see a Python list. It works with an ADDRESS and")
print("reads words from it. uctypes.addressof() gives us that address.")
print()

from array import array
from uctypes import addressof


@micropython.asm_thumb
def sum_array(r0, r1):
    # r0 = address of an array('i'), r1 = number of elements
    mov(r2, 0)                  # running total
    mov(r3, 0)                  # index

    label(SUM_LOOP)
    lsl(r4, r3, 2)              # index * 4  (each int is 4 bytes)
    add(r4, r0, r4)             # address of element
    ldr(r5, [r4, 0])            # load it
    add(r2, r2, r5)             # add to total
    add(r3, 1)
    cmp(r3, r1)
    blt(SUM_LOOP)

    mov(r0, r2)


data = array("i", [1, 2, 3, 4, 5, 6, 7, 8, 9, 10])
print("array      :", list(data))
print("sum_array  :", sum_array(addressof(data), len(data)), " (expect 55)")
print()
print("Note 'lsl(r4, r3, 2)' -- shifting left by 2 multiplies by 4, because")
print("each 32-bit integer occupies 4 bytes. Address arithmetic is manual.")

# =========================================================================
# PART 5 -- why bother
# =========================================================================
print()
print("=== PART 5: what it buys you ===")
print()

COUNT = 5000
big = array("i", [i for i in range(COUNT)])


def sum_python(a, n):
    total = 0
    for i in range(n):
        total += a[i]
    return total


def timeit(fn, *args):
    fn(*args)
    best = None
    for _ in range(5):
        s = rd()
        r = fn(*args)
        c = (rd() - s) & 0xFFFFFFFF
        if best is None or c < best:
            best = c
    return best, r


py_c, py_r = timeit(sum_python, big, COUNT)
asm_c, asm_r = timeit(sum_array, addressof(big), COUNT)

print("summing %d integers:" % COUNT)
print("  Python   : %8d cycles  -> %d" % (py_c, py_r))
print("  assembly : %8d cycles  -> %d" % (asm_c, asm_r))
print("  speedup  : %.0fx" % (py_c / asm_c))
print()
print("Same answer. The Python version spends almost all its time unboxing")
print("objects and dispatching bytecode. The assembly version does one")
print("load and one add per element, and nothing else.")

# =========================================================================
# The instruction set you actually need
# =========================================================================
print()
print("=== The dozen instructions that cover most of it ===")
print()
print("  mov(rd, imm)       put a constant in a register")
print("  mov(rd, rn)        copy a register")
print("  add(rd, rn, rm)    rd = rn + rm")
print("  add(rd, imm)       rd = rd + constant")
print("  sub(rd, rn, rm)    rd = rn - rm")
print("  mul(rd, rn)        rd = rd * rn")
print("  lsl(rd, rn, imm)   shift left (multiply by a power of 2)")
print("  lsr(rd, rn, imm)   shift right (divide by a power of 2)")
print("  ldr(rd, [rn, off]) load a 32-bit word from memory")
print("  str(rd, [rn, off]) store a 32-bit word to memory")
print("  cmp(rn, rm)        compare, setting flags")
print("  label(NAME)        mark a position")
print("  b(NAME)            jump")
print("  blt/ble/bgt/bne    jump if the comparison said so")
print()
print("That is genuinely most of it. Next lab adds the floating point set.")

Each part builds on the last, and the comments in the file explain the reasoning as you go. Run it, read it, then change something and run it again.

Predict before you measure

Echo offering a tip Wherever this lab reports a speedup, write your guess down before you run it. Every quantitative prediction made while building this course turned out to be optimistic — being wrong on paper is how you find out what the machine really does.

Troubleshooting

Symptom Likely cause Fix
unsupported Thumb instruction The assembler lacks that mnemonic Check Lab 28's probe; see Lab 33 for the workaround
Assembly returns nonsense Wrong argument order Arguments arrive in r0, r1, r2, r3
Results differ between runs No warm-up, or heap state Discard a warm-up; build objects before measuring (Labs 26, 32)
Variant looks slower than baseline Measurement artifact Re-run with everything allocated up front
MemoryError Too many variants alive gc.collect() between sections

Check Your Understanding

  1. What does this lab measure, and what does it deliberately exclude?
  2. Which result surprised you most against your prediction, and why?
  3. What would you change to make the effect larger?
  4. Where would this technique NOT be worth the complexity?

Onward

Echo celebrating You just wrote machine instructions and called them from Python. Next lab, the floating-point set.


Next: Lab 30 | Previous: Lab 28