Skip to content

Bring the hackathon branch up to date with main - #29

Merged
MatthiasHertelArm merged 22 commits into
hackathonfrom
hackathon-update-main
Oct 6, 2026
Merged

MatthiasHertelArm merged 22 commits into
hackathonfrom
hackathon-update-main

Conversation

@MatthiasHertelArm

Copy link
Copy Markdown
Contributor

Merges main into hackathon (a merge commit, so please merge this PR with Create a merge commit, not squash). It brings in #21, #23, #25, #26, #27 and #28:

  • ExecuTorch 1.5.1 (torch 2.14.0)
  • the NHWC model input
  • every Vela option of the mlops: node forwarded to Vela
  • .venv python in the Create AI layer task
  • CMSIS-Toolbox 2.15.0 without the CLANG -mfpu workaround
  • the FVP run directly as FVP_Corstone_SSE-320, with fvp.sh + .vscode/launch.json.mac as the macOS setup

Kept from hackathon: the DevKit-E8 target-type, the Ensemble packs and Vela system config, the Alif tasks, the DevKit-E8 CI job and the one-page README. CMSIS-Toolbox 2.15 detects the hardware (DevKit-E8) and simulator (SSE-320-U85) targets from the target-sets, so mlops: no longer names them. input-shape and calibration-samples are under model:.

Tested locally (toolbox 2.15.1, AC6): the regenerated AI layer is 8736 bytes, and both targets build. On the FVP the program prints the same logits as main and Test_result: PASS. No DevKit-E8 board was attached, so the hardware run is untested. CI covers both targets with AC6, GCC and CLANG.

After this is merged, hackathon-pico-faces (#22) gets the same update.

MatthiasHertel80 and others added 22 commits September 10, 2026 18:53
armclang rejects alignas after a GNU attribute ("'alignas' attribute
cannot be applied to types"); the attribute after the declarator is
accepted by AC6, GCC and Clang.
The FVP loaded the runner's libpython3.11 and died at start-up because
that build's standard library is not at its baked prefix; PYTHONHOME on
the FVP step points it at a complete Python 3.11. The flatbuffers pin
check used grep -q under pipefail, so grep closing the pipe early failed
the step on Linux and macOS.
--update-rte is no longer needed on a fresh checkout, the manual FVP run
uses the same shim as the Run and Debug buttons, the wheel pin that has
to match the pack is named, the project layout lists the task drop-in and
the shim, and the MLOps flow page names the six components the generated
layer selects. The Docker contrast is gone with the flow it referred to.
The region header claimed pack defaults that were this example's own
values; every Default line now states the SSE_320_BSP value. The layer
advertised a 768 kB heap while the scatter file reserves the 96 kB from
regions_SSE-320.h. The FVP configuration no longer refers to scripts and
tools this repository does not have.
…comments

The arm_embedded_module headers pointed at a LICENSE in the repository
root, which is Apache-2.0; src/LICENSE-ExecuTorch now carries the BSD
text. The generator's docstrings name every component it selects and why
the tensor extension is not among them, and the generated layer's Re-run
line names the argument the script needs.
CI no longer passes --update-rte (the RTE configuration is committed) and
keys the pack cache on the csolution so a pack bump refreshes it. The
.gitignore loses patterns for tools and directories this repository never
had and gains the FVP log the README and CI write. The shim's header shows
the arguments the launch config really passes and the Windows advice that
the README gives.
ExecuTorch 1.4.1 is on PyPI together with torch 2.13.0 and torchao
0.18.0, so the nightly index and requirements-executorch.txt go; every
pin lives in requirements.txt. The pack pin follows. The generated layer
now names the pack version cbuild setup resolved, and the generator reads
that version by the csolution's exact pin: the cbuild-pack.yml lock keeps
earlier resolutions, and an unversioned selector in the layer left the
previous pack listed next to the pinned one. Same 8832-byte program and
logits as with 1.4.0, verified on the FVP.
The Corstone-320 job becomes a toolchain matrix and runs the FVP for
each image; the LLVM embedded toolchain joins vcpkg-configuration.json.
Only AC6 was exercised before, which is how a component that does not
build with the LLVM toolchain could go unnoticed.
The Linux x86 build of ATfE 21.1.1 crashes in its register coalescer on
the quantize and dequantize kernels; 22.1.0 compiles them.
CMSIS-Toolbox 2.14.1 passes -mfpu=fpv5-sp-d16 to Clang for the
Cortex-M85 (and M55), a single-precision FPU on cores with a
double-precision FPU and MVE; LLVM 21.1.1 and 22.1.0 crash in the
register coalescer on the quantize and dequantize kernels with that
combination. A later -mfpu on the command line wins, so the csolution
adds -mfpu=fp-armv8-fullfp16-d16 for Clang; newer toolboxes already
leave the FPU to -mcpu.
create_ai_layer.py re-executes in .venv whenever the running interpreter
is not that venv; testing whether torch imports kept the export on a host
interpreter that happens to have torch. The Vela ini goes into the
compile spec as a relative path so the exported program does not embed
the checkout location. The pack-build recipe carries the CMake configure,
all generator arguments, the structural validator and a non-interactive
install; its checks now fail instead of echoing. Comments and the README
no longer attribute stdout to a UART, the flatbuffers pin to the
executorch wheel, or shutdown to uart0, and the adaptation section names
the runner's fixed input shape.
RTE/ is no longer partly ignored: the per-context RTE_Components.h and
Pre_Include_Global.h the toolbox writes are committed next to the layers'
configuration files.
… wheels

ExecuTorch 1.4.1 allows Python <3.15, but the tosa-tools 2026.5.0 it pins
for the Arm backend only ships wheels for CPython 3.10 to 3.13. On 3.14
pip fails with "No matching distribution found for tosa-tools==2026.5.0"
long after the venv was created. Check the bound up front and document it.
setup_venv.py and setup_venv.sh refuse 3.14 since tosa-tools 2026.5.0 has
no wheels for it; the Windows wrapper's default range for uv and the prompt
of the uv setup task still said <3.15.
create_ai_layer.py picked --accelerator-config, --system-config and
--memory-mode out of vela.options of the cbuild-mlops.yml with a regular
expression and dropped the rest, so `vela: misc:` in the csolution had no
effect and nothing said so. The options are now tokenized; the three that
are arguments of EthosUCompileSpec go there and every other one reaches
Vela as an extra flag. An option given twice, and --config,
--output-format and --output-dir, which the export sets itself, end the
script with a message.

Without misc: the exported program is byte for byte the one in ai_layer/.
The PyTorch::ExecuTorch pack 1.5.1 is in the public pack index and
executorch 1.5.1 is on PyPI. It targets torch 2.14.0 (torch_pin.py at the
v1.5.1 tag) and keeps torchao 0.18, ethos-u-vela 5.1.0 and tosa-tools
2026.5.0, so requirements-arm-tosa.txt is unchanged. The pack pins in the
csolution and the cproject follow, and the generated layer and the RTE
context headers name the new pack version.

The exported program is 8784 bytes instead of 8832.
TinyCNN took its image as NCHW. The Ethos-U computes on NHWC feature
maps, so the command stream began with an operation that only reorders
the input. The model now takes NHWC, the layout of a camera frame, and
permutes it to NCHW itself; Vela folds that into the first convolution.

The layout conversion is gone from the command stream, the program is
8720 bytes instead of 8784, and the Ethos-U scratch 5376 bytes instead
of 8704. The runner declares its input as 1x16x16x3; its ramp now fills
the image in that order, so the logits differ from the ones before.
….0 (#27)

The target-set's model: is FVP_Corstone_SSE-320 again, so Run and Debug work on Windows and Linux with the launch.json the extension generates; the Docker shim and its launch configuration (now .vscode/launch.json.mac) are the macOS setup, described in the README. CMSIS-Toolbox 2.15.0 no longer passes -mfpu=fpv5-sp-d16 to Clang for the Cortex-M85, so the CLANG misc: workaround goes; vcpkg pins the toolbox at 2.15.0.
From the preview branch, now that CMSIS-Toolbox 2.15.0 is released: the toolbox detects the simulator target from the Arm-FVP target-set, so simulator: goes, and passes extra keys under model: through to *.cbuild-mlops.yml. The NHWC input shape and the number of calibration samples move there; model_params() in create_ai_layer.py hands them to model/model.py. The exported program is unchanged.
MD schema compliance fix.
…WC input

Brings the hackathon branch up to date with main (#21, #23, #25, #26,
#27, #28): ExecuTorch 1.5.1 with torch 2.14.0, the NHWC model input,
every Vela option of the mlops: node passed to Vela, the .venv python in
the Create AI layer task, CMSIS-Toolbox 2.15.0 without the CLANG -mfpu
workaround, and the FVP run directly with fvp.sh plus launch.json.mac as
the macOS setup.

The hackathon parts stay: the DevKit-E8 target-type, the Ensemble packs,
the Ensemble Vela system config, the Alif tasks, the DevKit-E8 CI job and
the one-page README. CMSIS-Toolbox 2.15 detects both the hardware and the
simulator target from the target-sets, so neither is named in mlops:
any more; input-shape and calibration-samples move under model:.

The regenerated AI layer is 8736 bytes and prints the same logits as main
on the FVP; the README, example.md and mlops-flow.md follow.
@MatthiasHertelArm
MatthiasHertelArm merged commit 2ac7d2f into hackathon Oct 6, 2026
9 checks passed
@MatthiasHertelArm
MatthiasHertelArm deleted the hackathon-update-main branch October 6, 2026 14:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants