Making stinkarm stink less, or more?

Hacker News Top Tools

Summary

An update on improving a minimal ARMv7 emulator, focusing on optimizing memory translation from complex B-tree mapping to a simplified 4 GiB slab approach for efficiency.

No content available
Original Article
View Cached Full Text

Cached at: 08/21/26, 04:30 PM

# Making stinkarm stink way less, or more? Source: [https://xnacly.me/posts/2026/making-stinkarm-stink-less/](https://xnacly.me/posts/2026/making-stinkarm-stink-less/) ![acorn, armv7-a and stinkarm](https://xnacly.me/stinkarm/stinkarm.png)[1](https://xnacly.me/posts/2026/making-stinkarm-stink-less/#fn:1) Its been a while, but about half a year ago I wrote an article about implementing a userspace armv7 emulator from scratch, meaning I implemented: ARMASM ``` 1 .section .rodata 2msg: 3 .asciz "Hello, world!\n" 4 5 .section .text 6 .global _start 7_start: 8 ldr r0, =1 9 ldr r1, =msg 10 mov r2, #14 11 mov r7, #4 12 svc #0 13 14 mov r0, #0 15 mov r7, #1 16 svc #0 ``` Or as a list: - elf\(32\) parsing, validation and interpretation - decoding of a very small subset of armv7 instructions \(only 3\) - executing said instructions, even conditional ones 🤓 - translating memory access from the guest into the host - syscall forwarding \(from armv7 to x86\) - syscall sandboxing \(only a restricted syscall subset\) and denying syscall execution Do read[Building a Minimal Viable Armv7 Emulator from Scratch](https://xnacly.me/posts/2025/building-a-minimal-viable-armv7-emulator/), since this post doesnt go as deep into detail as the previous one \(It’s my first article in 3 months I had enough motivation for writing :O\)\. This is partially an update, partially my toughts on decoding and emulating armv7\-a and also a bit of a devlog\. ## Overly complex host to guest mem translation On the first article,[~aengelke on lobste\.rs](https://lobste.rs/s/bv3570/building_minimal_viable_armv7_emulator), had some comments, the one resonating the most was: > \[…\] The Mem indirection seems pretty inefficient\. When emulating 32\-bit platforms on a 64\-bit system, just mmap a 4 GiB region, the translation then becomes a single addition\. Otherwise, having a small hash table of recently translated address regions can avoid more expensive searches – memory accesses have a very high locality\. The number of mappings is usually small, so binary search over a sorted array is simpler than a B\-tree\. \[…\] So now i figured, why not improve on my implementation a bit, first with replacing the complex allocation region based tracking with just allocating a 4gig slab in memory for the guest, mapping the process regions there and handing out pointers into that region to the guest\. So, previously the memory translation worked as follows: 1. One takes a binary tree map of a guest starting addr to its host segment RUST ``` 1struct MappedSegment { 2 host_ptr: *mut u8, 3 len: u32, 4} 5 6pub struct Mem { 7 maps: BTreeMap<u32, MappedSegment>, 8} ``` 2. On ask for a region handout, specifically on mapping ELF segments with a starting addr,`map\_region`is called: RUST ``` 1// in stinkarm::elf::pheader::Pheader::map: 2 3// record mapping in guest memory table, so CPU can translate guest vaddr to host pointer 4guest_mem.map_region(self.vaddr, len, segment_ptr); 5 6// in stinkarm::mem::Mem: 7 8pub fn map_region(&mut self, guest_addr: u32, len: u32, host_ptr: *mut u8) { 9 self.maps 10 .insert(guest_addr, MappedSegment { host_ptr, len }); 11} ``` 3. Since the cpu needs to fetch an instruction, there is`read\_u32`, calling`translate`: RUST ``` 1/// translate a guest addr to a host addr we can write and read from 2pub fn translate(&self, guest_addr: u32) -> Option<*mut u8> { 3 // Find the greatest key <= guest_addr. 4 let (&base, seg) = self.maps.range(..=guest_addr).next_back()?; 5 if guest_addr < base.wrapping_add(seg.len) { 6 let offset = guest_addr.wrapping_sub(base); 7 Some(unsafe { seg.host_ptr.add(offset as usize) }) 8 } else { 9 None 10 } 11} 12 13pub fn read_u32(&self, guest_addr: u32) -> Option<u32> { 14 let ptr = self.translate(guest_addr)?; 15 unsafe { Some(u32::from_le(*(ptr as *const u32))) } 16} 17 18 19// in stinkarm::cpu::Cpu: 20 21pub fn step(&mut self) -> Result<bool, err::Err> { 22 let Some(word) = self.mem.read_u32(self.pc()) else { 23 return Ok(false); 24 }; 25 26 // [...] 27} ``` Of course this totally unnecessary work, we dont need to keep track of every mapping/allocation/region by walking their ranges, we only need to make sure the R/W interaction request is within bounds\. Thus the new implementation is: 1. One takes a pointer and a size: RUST ``` 1pub struct Mem { 2 ptr: NonNull<u8>, 3 len: usize, 4} ``` 2. When asked to map ELF segments,`stinkarm::mem::Mem::map\_region`is called: RUST ``` 1// in stinkarm::elf::pheader::Pheader::map: 2guest_mem.map_region(self.vaddr, file_slice)?; 3 4// in stinkarm::mem::Mem: 5 6pub fn map_region(&mut self, guest_addr: u32, data: &[u8]) -> Result<(), String> { 7 let dst = self 8 .get_slice_mut(guest_addr, data.len()) 9 .ok_or_else(|| format!("guest region out of bounds at {guest_addr:#010x}"))?; 10 dst.copy_from_slice(data); 11 Ok(()) 12} ``` 3. When cpu requests a dword for decoding, it does so by invoking`stinkarm::mem::Mem::read32`, just as before, only this time with bounds checks: RUST ``` 1pub fn read_u32(&self, guest_addr: u32) -> Option<u32> { 2 let bytes = self.get_slice(guest_addr, 4)?; 3 Some(u32::from_le_bytes(bytes.try_into().unwrap())) 4} 5 6fn get_slice(&self, guest_addr: u32, len: usize) -> Option<&[u8]> { 7 if !self.in_bounds(guest_addr, len) { 8 return None; 9 } 10 11 Some(unsafe { std::slice::from_raw_parts(self.ptr.as_ptr().add(guest_addr as usize), len) }) 12} ``` ## Hardening the existing implementation I also noticed I have a lot of stuff that \(even with the small surface of just the`write\.2`and`exit\.2`syscalls,`ldr`,`mov`and`svc`\) could enable translating untrusted guest adresses into host mem access\. Preventing this via checking the validity of the address passed to`write\.2`, we do this while translating guest addresses to host memory space in`stinkarm::mem::Mem`with a`in\_bounds`call inside the`translate\_range`call: RUST ``` 1const NULL_PAGE_SIZE: u32 = 0x1000; 2 3impl Mem { 4 fn in_bounds(&self, guest_addr: u32, len: usize) -> bool { 5 if guest_addr < NULL_PAGE_SIZE { 6 return false; 7 } 8 9 let start = guest_addr as usize; 10 let Some(end) = start.checked_add(len) else { 11 return false; 12 }; 13 14 end <= self.len 15 } 16 17 pub fn translate_range(&self, guest_addr: u32, len: usize) -> Option<*mut u8> { 18 if !self.in_bounds(guest_addr, len) { 19 return None; 20 } 21 22 Some(self.ptr.as_ptr().wrapping_add(guest_addr as usize)) 23 } 24} ``` I added multiple tests for making sure I correctly catch writing a null pointer, writing out of guest memory and loading elf segments at 0x0: ARMASM ``` 1 .section .rodata 2msg: 3 .ascii "ignored" 4 5 .section .text 6 .global _start 7_start: 8 mov r0, #1 9 mov r1, #0 10 mov r2, #7 11 mov r7, #4 12 svc #0 13 14 mov r0, #0 15 mov r7, #1 16 svc #0 ``` ARMASM ``` 1 .section .rodata 2msg: 3 .ascii "ignored" 4 5 .section .text 6 .global _start 7_start: 8 mov r0, #1 9 ldr r1, =0x08000000 10 mov r2, #7 11 mov r7, #4 12 svc #0 13 14 mov r0, #0 15 mov r7, #1 16 svc #0 ``` ## To DSL or not Previously I hardcoded every opcode and its fields to decode them into a rust representation, now only the opcode is subject to decoding\. This is achived with a nice looking compiletime constant list of patterns: RUST ``` 1const DECODE_RULES: &[ArmRule] = &[ 2 arm_rule!(Svc { 3 bits(27..24 = 0b1111), 4 }), 5 arm_rule!(Branch { 6 bits(27..25 = 0b101), 7 }), 8 // LDR literal: `ldr Rt, [pc, #imm12]`. 9 arm_rule!(LdrLiteral { 10 bits(27..26 = 0b01), // load/store class 11 bit(24 = 1), // P: pre-indexed address 12 bit(23 = 1), // U: add positive offset 13 bit(22 = 0), // B: word transfer, not byte 14 bit(21 = 0), // W: no writeback 15 bit(20 = 1), // L: load, not store 16 bits(19..16 = 15), // Rn: base register is pc/r15 17 }), 18 // MOV immediate: data-processing immediate with opcode 1101. 19 arm_rule!(MovImm { 20 bits(27..25 = 0b001), 21 bits(24..21 = Op::Mov as u32), 22 }), 23]; ``` > If youre interested in ARMv7 instruction encoding I can recommend the[ARM® Architecture Reference Manual ARMv7\-A and ARMv7\-R edition](https://documentation-service.arm.com/static/5f8daeb7f86e16515cdb8c4e) The macro itself builds a bit pattern that can then be used with a simple AND bit instruction to detect: RUST ``` 1macro_rules! arm_rule { 2 ($kind:ident { $($field:ident($($args:tt)*)),* $(,)? }) => { 3 ArmRule { 4 kind: InstructionKind::$kind, 5 mask: 0 $(| arm_mask!($field($($args)*)))*, 6 value: 0 $(| arm_value!($field($($args)*)))*, 7 } 8 }; 9} 10 11macro_rules! arm_mask { 12 (bit($bit:literal = $value:expr)) => { 13 1u32 << $bit 14 }; 15 (bits($high:literal .. $low:literal = $value:expr)) => { 16 ((1u32 << ($high - $low + 1)) - 1) << $low 17 }; 18} 19 20macro_rules! arm_value { 21 (bit($bit:literal = $value:expr)) => { 22 ($value as u32) << $bit 23 }; 24 (bits($high:literal .. $low:literal = $value:expr)) => { 25 ($value as u32) << $low 26 }; 27} ``` `MovImm`’s definition does therefore produce \(`Op::mov`is defined as`0b1101`\): RUST ``` 1ArmRule { 2 kind: InstructionKind::MovImm, 3 mask: 0 4 | ((1u32 << (27 - 25 + 1)) - 1) << 25 5 | ((1u32 << (24 - 21 + 1)) - 1) << 21, 6 value: 0 7 | (0b001 as u32) << 25 8 | ((Op::Mov as u32) as u32) << 21, 9} 10 11 12impl ArmRule { 13 fn matches(&self, word: u32) -> bool { 14 (word & self.mask) == self.value 15 } 16} ``` All rules are then iterated for each 32bit word and decoded: RUST ``` 1pub fn decode_word(word: u32) -> Decoded { 2 let cond = bits(word, 31, 28) as u8; 3 let kind = DECODE_RULES 4 .iter() 5 .find(|rule| rule.matches(word)) 6 .map(|rule| rule.kind) 7 .unwrap_or(InstructionKind::Unknown); 8 9 Decoded { 10 cond, 11 kind, 12 raw: word, 13 } 14} ``` I know a trie or something would probably be faster, but this is nice to read, understandable and easy to maintain\. ## Full decoding only on demand Previous to the instruction DSL, I decoded all necessary values for all instructions at all times, meaning short instructions would decode even if they didnt match, just as a sideeffect of attempting to figure out what instruction\. The DSL allows only decoding the op code and letting the cpu decode only what it needs via`decoder::\{decode\_word,bit,bits\}`and`decoder::\{sign\_extend,rotated\_imm\}`where bit and bits enable partial access into the word, and decode\_word returns the op code, the condition and the raw word itself for further processing in the cpu emulation: RUST ``` 1/// fetch-decode-execute step, will only return false on exit svc 2pub fn step(&mut self) -> Result<bool, err::Err> { 3 // [...] fetch word 4 5 let Decoded { kind, cond, raw } = decoder::decode_word(word); 6 7 // [...] 8 9 match kind { 10 InstructionKind::MovImm => { 11 let rd = decoder::bits(raw, 15, 12) as usize; 12 let imm12 = decoder::bits(raw, 11, 0); 13 // [...] 14 } 15 // [...] 16 InstructionKind::LdrLiteral => { 17 let rd = decoder::bits(raw, 15, 12) as usize; 18 let imm12 = decoder::bits(raw, 11, 0); 19 20 // [...] 21 } 22 } 23} ``` Bits and bit access is obvious, sign\_extend and rotated\_imm maybe less so: RUST ``` 1pub fn bits(word: u32, high: u8, low: u8) -> u32 { 2 debug_assert!(high < 32); 3 debug_assert!(low <= high); 4 let width = high - low + 1; 5 (word >> low) & ((1 << width) - 1) 6} 7 8pub fn bit(word: u32, bit: u8) -> bool { 9 bits(word, bit, bit) != 0 10} 11 12pub fn sign_extend(value: u32, bits: u32) -> i32 { 13 debug_assert!((1..=32).contains(&bits)); 14 15 let shift = 32 - bits; 16 ((value << shift) as i32) >> shift 17} 18 19pub fn rotated_imm(imm12: u32) -> u32 { 20 let rotate = ((imm12 >> 8) & 0b1111) * 2; 21 (imm12 & 0xff).rotate_right(rotate) 22} ``` ## Supporting B and BL B and BL are the unconditial branching instructions of the ARMv7 isa: - Branching unconditionally: ARMASM ``` 1.text 2 .global _start 3_start: 4 mov r0, #0 5 b 1f 6 mov r0, #1 @ must NOT execute 71: 8 mov r7, #1 9 svc #0 ``` - Branching unconditionally with link: ARMASM ``` 1.text 2 .global _start 3_start: 4 mov r0, #0 5 bl foo 6 mov r7, #1 7 svc #0 8foo: 9 mov r0, #42 10 mov r7, #1 11 svc #0 ``` B and BL encode: 1. Conditionals, see[ARMv7 Condition code suffixes](https://support.arm.com/documentation/den0042/0100/Unified-Assembly-Language-Instructions/Instruction-set-basics/Conditional-execution?lang=en#md260-conditional-execution__tbl_cond_code_suffixes) 2. Instruction group \(`101`\) 3. Wheter or not to branch with link \(`L`\) 4. Target \(imm24\) In bits: TEXT ``` 1 31 30 29 28 27 26 25 24 23 .. 0 2|cond |1 0 1 |L | imm24 | ``` Meaning, its fairly easy to implement, see below\. L instructs the emulator to save the return addr to the**L**ink**R**egister \(LR\) and otherwise we just decode the imm24, shift it 2 to the left and then sign extend it to 32 bit, add it to the program counter and thats it: RUST ``` 1InstructionKind::Branch => { 2 let l = decoder::bit(raw, 24); 3 // BL 4 if l { 5 // save return addr to LR (next addr though) 6 self.r[14] = self.instr_addr().wrapping_add(4); 7 } 8 9 let imm24 = decoder::bits(raw, 23, 0); 10 let imm26 = imm24 << 2; 11 let imm32 = decoder::sign_extend(imm26, 26); 12 13 self.r[15] = self.arm_pc().wrapping_add(imm32 as u32); 14} ``` ## Testing “frame\(work\)” To test all the hardening stuff and every new instructions I intent to support, i added a bit of tooling, specifically srun, its a small stinkarm wrapper to build, link and execute assembly or c files: TEXT ``` 1Build, link, and execute an ARM assembly or C file with stinkarm 2 3Usage: srun [OPTIONS] <INPUT> [-- <EMULATOR_ARGS>...] 4 5Arguments: 6 <INPUT> ARM assembly or C file to run 7 [EMULATOR_ARGS]... Extra arguments passed to stinkarm before the generated ELF path 8 9Options: 10 --text-addr <TEXT_ADDR> Guest address used as the linker text address [default: 0x8000] 11 --out-dir <OUT_DIR> Directory for generated object and ELF files [default: target/srun] 12 --dump-asm Print the linked ARM disassembly before running the emulator 13 -h, --help Print help ``` For instance previously I had to first assemble the`examples/branch\.S`file, then invoke stinkarm, now i can just: SHELL ``` 1cargo run --bin srun 2 # arguments for srun 3 \ -- examples/helloWorld.S 4 # arguments for stink arm, log instructions and syscalls 5 \ -- -linstructions -lsyscalls ``` TEXT ``` 1[ 0.538ms] MovImm 1110 E3A00001 2[ 0.543ms] LdrLiteral 1110 E59F1014 3[ 0.545ms] MovImm 1110 E3A0200E 4[ 0.548ms] MovImm 1110 E3A07004 5[ 0.551ms] Svc 1110 EF000000 665174 write(fd=1, buf=0x8024, len=14) [sandbox] 7Hello, world! 8=14 9[ 0.567ms] MovImm 1110 E3A00000 10[ 0.570ms] MovImm 1110 E3A07001 11[ 0.573ms] Svc 1110 EF000000 1265174 exit(code=0) [sandbox] 13=0 ``` So yeah, thats it, now please enjoy me bashing in a claude server rack: ![claude](https://xnacly.me/stinkarm/claude-ze-fucker.png)

Similar Articles

Pushing the limits of RISC-V emulation

Hacker News Top

This blog post explores how to make RISC-V execution faster on non-RISC-V machines using an ahead-of-time recompiler that connects basic blocks with tail calls and leverages Clang's preserve_none calling convention, as part of the OpenVM project at Axiom.

PS3 emulation is fast on ARM now

Lobsters Hottest

RPCS3's ARM port now runs 60% faster and draws 25% less power after fixing a busy-wait timer bug, replacing x86 pause with ARM ISB, and reworking LLVM code generation. The gains come from low-level ARM optimizations driven by a cheap Android handheld test device.

Optimizing LLVM's bump allocator

Lobsters Hottest

This blog post details three recent optimizations to LLVM's BumpPtrAllocator, reducing fast-path overhead by removing redundant alignment, null pointer checks, and per-allocation accounting, resulting in improved performance for Clang, lld, and other LLVM components.

Windows stack limit checking retrospective, follow-up

The Old New Thing (Raymond Chen)

Raymond Chen follows up on his previous article about stack limit checking on ARM64, addressing a detail about the unconventional use of the x15 register in stack probe functions and comparing register usage across multiple architectures.