GCC 14+ defaults to gnu23, where glibc string.h defines memchr
and memcpy as _Generic macros that clash with wolfBoot's own
declarations in tests that #include a .c file (unit-string).
gnu17 matches the CI toolchain default.
keygen_die() (F-9767) calls wc_FreeRng(); the test mocks
keygen.c's wolfCrypt dependencies one by one and was missing
this one, breaking the unit_tests and test_external_libs jobs.
Secure-mode worlds (TZ_PSA/PKCS11/FWTPM/WOLFHSM) process private keys
in software, so a USE_FAST_MATH build of one of them must keep the
timing-resistant TFM path instead of falling back to WC_NO_HARDEN.
The old condition only covered software DICE; FIPS=1 already forces
SPMATH (options.mk M10) and all CMake presets set SPMATH, so no
shipped config changes - this closes the integrator Makefile gap.
WOLFBOOT_DICE_HW stays verify-only (requires WOLFCRYPT_TZ_PSA).
buffer_is_all_value() early-exited on the first byte differing from
the target, so the loop trip count leaked the length of the leading
0xFF/0x00 run of the UDS, the DICE root secret read in
hal_uds_derive_key().
Replace the early exit with a volatile |= accumulator over the full
buffer, matching the constant-time compare pattern already used by
image_CT_compare() and the other secret comparisons in the tree.
Verified: gcc -S -O2 shows a straight-line loop body with only the
data-independent i < len branch; arm-none-eabi-gcc (stm32h5 preset)
compiles the file clean.
Reported by Fenrir.
security_command_passphrase() only wiped the static DMA buffer on the
synchronous path; in async mode (ata_security_erase_unit) the
passphrase stayed resident in the file-static buffer for the rest of
the boot, since ata_cmd_complete_async() never scrubbed it.
Add a scrub_buffer flag to the async state, set it when the command
goes in flight (or wipe immediately if it never started), and scrub in
ata_cmd_complete_async() on both the success and task-file-error exits,
once the HBA has retired the command.
Covered by two new unit tests: scrub after async completion and after
an async port error; the in-flight state is asserted unscrubbed.
Reported by Fenrir.
ARM_TEE_PS_SET on an existing object XMEMCPY'd the new value over the
old one without clearing the buffer first, so a SET that stores less
data (or zero data) left the tail of the previous object readable
through ARM_TEE_PS_GET.
Zero the full data area with wc_ForceZero after every validation check
passes and before the copy, matching the DELETE path.
Reported by Fenrir.
The TPM NV auth copy and the HMAC session state lived on the
stack until process teardown. Zeroize both at the common exit
label after the TPM device is unloaded, on success and error
paths alike.
keygen_ed25519/ed448/lms/xmss/ml_dsa have the same bypass as
the RSA/ECC helpers: exit() from their cleanup labels skips
main()'s wc_FreeRng(). Route their failure exits through
keygen_die() as well.
keygen_rsa/keygen_ecc exit() from their own cleanup labels,
bypassing main()'s wc_FreeRng() and leaving DRBG state resident.
Add keygen_die() that frees and zeroes the RNG before exit and
use it in the two helpers' failure paths.
The error path returned before the common wc_FreeRng(), leaving
DRBG state on the stack. Route it through the shared cleanup and
guard the sign call so hash_type/mgf are only used when set.
The F-12922 test checks fallback to the lower version, but the
TCase was still named "rollback denied" from before the fix,
describing the opposite behavior.
The F-12879 fix publishes SHM status with plain C stores before
IPC_TASKS_SEND; without a compiler memory barrier an optimized
build can reorder the stores past the dsb asm. Matches the
existing clobber in imx95_m7.h and stm32g4.h.
ext_flash_erase() now checks NAND_STATUS_FAIL | NAND_STATUS_WPS
(F-12881), but the sed extraction rule only pulled the command
defines, so the unit_tests CI jobs failed to compile the test.
pkcs11_crypto_init() tore down the session on failure but left the
pre-populated pkcs11_pin credential in retained bootloader memory
after the handoff. Wipe it in the failure path, same as the deinit
path (F-12114). Regression test in unit-pkcs11-pin-zeroize: login
rejected -> init fails -> every pin byte is zero.
hal_shm_status_set() triggered the IPC event before storing the
magic/status fields, so the peer core could wake on the event, read
stale fields, and wait for a second event that never comes. Store
the fields first, order them with a DSB, then signal; the receiver
pairs it with a DSB after clearing the event, before re-reading the
fields. Not verified on real nRF5340 hardware.
The write/erase loops discarded the JEDEC status byte from MDR and
hal_flash_command ignored the PAR (uncorrectable ECC) and FCT (FCM
timeout) bits, so failed programs, erases and uncorrectable reads
all reported success. Check P/WPS in the status byte and stop the
loop, and fail the command on PAR/FCT. The small-page FCM sequence
now ends with the status command (CM3+RSW) so MDR holds a valid
status like the large-page path; not verified on real P1021
hardware. config_io_pin now uses a single masked store for CPDIR/
CPPAR (F-12880): the clear-then-set pair could drop a concurrent
update to another pin in the same register.
F-12880, F-12881, F-12882
The version gate in the retry loop panicked whenever the selected
candidate failed verification and the fallback partition carried a
lower version, even though that image is valid and signed. A single
corrupted higher-version image bricked the device instead of falling
back. The gate only ever fired on the failure path (candidate
selection already prefers the higher version), so remove it; the
TESTING-state anti-rollback check in wolfBoot_dualboot_candidate()
still guards the confirmed-update path.
wolfBoot_start() zeroed os_image once, before the retry loop, so the
fallback iteration kept the stale img->hdr of the failed partition and
wolfBoot_open_image_address() (which only adopts load_address when
hdr is NULL) re-verified the wrong image. The external header cache
had the same problem: it kept the first image opened.
memset the struct and invalidate the EXT_FLASH header cache at the
top of each loop iteration, per the documented precondition of
wolfBoot_open_image_address().
Add unit-update-ram-nofixed-noramboot: the non-fixed-partition,
non-RAMBOOT (XIP) layout where load_address varies per retry, with
fallback tests both directions plus candidate-selection checks.
The mock abandoned a faulting erase/program whole, so half-written sectors
went untested; it can now tear one part-way and the sweep runs every crash
point at 1/4, 1/2 and 3/4. An empty read must also leave the node in place
with size 8, so a lost or torn header no longer passes as an empty rewrite.
The crash == ops iteration injects nothing, so it must read back new_p; the
empty/absent branch was letting a silently lost complete rewrite pass. Also
assert the uninjected baseline write lands before measuring against it.
Under MOCK_STALE_CACHE the shadow is the flash array and vault_base only the
CPU's view, so restoring vault_base alone let the power cycle copy the last
iteration's result back: no crash case started from the old generation.
vault_obj_read() now separates "object absent" from a failed read, so the
power-fail test rejects a corrupted vault instead of accepting any negative.
hal_cache_invalidate() shrank 120 -> 44 bytes; the 48 it still costs every
STM32F407 build is folded into the test-size-all limits.
The weak no-op sat inside NVM_FLASH_WRITEONCE, so pkcs11_store.c failed to
link on the nrf5340/nrf54l TrustZone configs. STM32F4 and STM32G4 enable the
flash instruction/data caches and never reset them, so give them a real one.
The store commits sectors with hal_flash_erase()/hal_flash_write() and
then reads them back through the memory map - sector_ptr(),
cache_get_sector()'s refill, and the raw magic reads in check_vault().
On a part that caches flash reads (STM32 ICACHE) those reads can return
pre-erase bytes. check_vault() is the worst case: a stale magic there
does not merely read wrong, it triggers restore_backup() or a full vault
re-initialisation, losing the token.
Adds power-fail injection to the flash mock and a test that cuts power at
every flash operation of a rewrite window, asserting the object always reads
back as the whole old payload, the whole new payload, or empty. Fails at
op 3 without this fix, passes with it.
unit-pkcs11_store 10/10; full unit-tests suite green.
Three fixes from the 2026-08-31 Fenrir PR review round:
1. cache_get_sector() LRU eviction could pick the header sector (offset 0)
as the victim, committing it to flash while the batch's payload sectors
were still only in RAM - a mixed pre/post-batch state that breaks the
header-last atomic commit point cache_flush_all() relies on. The header
is now exempt from victim selection; if it is the only cached sector,
flush the whole batch instead (header-last is then trivial).
2. cache_flush_all() and LRU eviction released a slot by clearing .sector
without wiping the buffer, so private-key bytes staged by Store_Write
lingered in secure-world SRAM until the slot was next reused. The single
staging buffer this cache replaced self-cleaned (the header sector
overwrote it on every size update); per-sector slots do not. A new
cache_release() wc_ForceZero()s the buffer before freeing the slot.
3. Store_Open in write mode set the size to 8 (truncation) in the cache
only; with batched commits the payload sectors are flushed before the
header, so a power loss during the erase/rewrite left the old size over
a partly erased payload. The truncated header is now committed to flash
before erase_object_payload(), so the empty state is the crash fallback.
Addresses PR #873 review comments (wolfSSL-Fenrir-bot,
src/pkcs11_store.c:301, :333, :670, 2026-08-31).
Verification: tools/unit-tests unit-pkcs11_store 9/9 pass (incl.
test_concurrent_reader_sees_pending_writes,
test_shorter_overwrite_erases_residual_key_material,
test_interleaved_write_windows_both_persist).
test_concurrent_reader_sees_pending_writes only ever read from flash:
the reader's Store_Open calls check_vault(), which flushes the sector
cache, so the sector_ptr() read path in Store_Read was never exercised
and the test passed identically against the pre-PR memcpy.
Write more on the still-open writer after the reader is open. That
batch lands only in the sector cache, so the reader can only see it
through the cached read path and the live header size; a flash-only or
snapshot-size read returns EOF here. Verified: the new assertion fails
against the pre-fix store (ret == 0) and passes with the live-size fix.
Addresses PR #873 review comments (wolfSSL-Fenrir-bot,
tools/unit-tests/unit-pkcs11_store.c:587, both near-duplicate findings).
Store_Read and Store_Write used handle->size, a snapshot taken at
Store_Open. The payload path reads through the sector cache, so once
another window's batch (e.g. a write-open truncation) sat pending in
the cache, the window saw live erased data under a stale size and
returned 0xFF bytes past the true end instead of EOF. Pre-PR the size
was read live from the flash header on every call, so the PR regressed
that case.
Read the size from the same (possibly cached) header sector the payload
comes from, via store_live_size(), so size and data share one source of
truth. Drop the now-dead handle->size snapshot; update_store_size()
only writes the cached header node.
Addresses PR #873 review comment (wolfSSL-Fenrir-bot,
src/pkcs11_store.c:711).
check_vault() dropped the shared sector cache on every vault
validation, silently losing the pending writes of any still-open
window when another handle was opened or an object removed
(MAX_OPEN_STORES allows 16). Flush instead - the atomic header-last
commit - so an in-flight batch only gets an earlier commit point;
its data is never discarded.
wolfPKCS11_Store_Read() now reads through sector_ptr() like every
other read in the file, so a sector still in the cache can never be
read stale against a live size.
Add unit tests covering the interleaved-window data loss and a
concurrent reader observing a pending write; both fail without the
check_vault fix.
Every wolfPKCS11 field write flushed the payload sector and the header
sector to flash (2 erases + 2 programs of a full sector each), and the
token store re-serializes all objects per C_CreateObject/C_DestroyObject,
so those calls cost hundreds of sector erases and tens of seconds on
flash with slow erase times.
Cache modified sectors in RAM and commit them together when the store
window closes:
- sector cache sized to the worst-case span of one object plus the
header sector (WOLFBOOT_PKCS11_STORE_CACHE_SECTORS), LRU eviction
when exceeded
- header sector commits last, so a committed header is the atomic
commit point of the batch: power failure during a flush leaves the
flash in either the pre-batch or the post-batch state
- per-commit backup sector write preserved, keeping recovery of the
sector in flight at failure time
- delete_object commits on return (durability contract, unit-tested)
- nodes table, bitmap, payload ids and the live object size
(handle->size) are read from the cache when the sector is dirty
Measured on an STM32H5 with 8KB sectors, wolfPKCS11 in the secure
world: C_CreateObject 1.5s -> 0.15s, C_DestroyObject 1.3s -> 0.12s,
456 -> 40 sector erases per create, and the count no longer scales
with the number of objects in the token.
PKCS11_STORE_STATS (off by default) adds flash-activity counters and a
test-app bench to quantify store traffic: make PKCS11_STORE_STATS=1.
The new unaligned tests passed against the pre-fix HAL (identical
final bytes), so the alignment fix had no regression coverage. Mock
hal_flash_wait_complete now diffs the flash per program window and
asserts the changed bytes fit in one aligned unit; the 20-byte
unaligned test goes red on the pre-fix HAL (l5: bytes 4-11 across two
8-byte units, u5: bytes 4-19 across two 16-byte units).
hal/stm32wb.c: dedicated RCC_CFGR_SWS_{MSI,MASK} macros for the
clock-switch confirmation wait; SW/SWS encodings verified identical
in RM0434 6.4.3 and the STM32WB55 SVD.
unit-stm32u5-write.c: START_TEST brace on the next line, matching
the file and the unit-suite convention.
An unaligned starting address split the four word stores across two
16-byte program units, leaving partial quad-words that set
FLASH_SR_WDW and hang the wait for completion. Align the destination
down to the unit, read-modify-write the whole unit, and store through
the aligned pointer. The unit test gains unaligned-start cases.