Skip to content

Add runnable Windows local AI development setup - #104

Open
Michael Von Hippel (Kixantrix) wants to merge 23 commits into
microsoft:mainfrom
Kixantrix:mihippel-microsoft-windows-ai-setup-workloads
Open

Michael Von Hippel (Kixantrix) wants to merge 23 commits into
microsoft:mainfrom
Kixantrix:mihippel-microsoft-windows-ai-setup-workloads

Conversation

@Kixantrix

@Kixantrix Michael Von Hippel (Kixantrix) commented Sep 17, 2026 •

Copy link
Copy Markdown

Deliverable

Add a runnable Windows local-AI development scenario plus independent vendor/runtime flows:

  1. detect Windows CPU architecture and accelerators;
  2. install the hardware-appropriate contained PyTorch backend;
  3. install compatible Triton only where a Windows package is published;
  4. execute a tensor and minimal neural-network forward pass on the intended device;
  5. optionally install one local-model runtime: llama.cpp, Ollama, or Foundry Local.

This supersedes closed #98. That PR targeted #93's feature branch; #93 has since merged into main and its branch was deleted. This branch was rebuilt onto current main and preserves the final AI implementation and physical validation.

Non-goals

  • This is not a replacement for PyPI, Conda, or application dependency management.
  • It does not install every vendor SDK or every local-model runtime.
  • ARM64 is a CPU architecture, not a GPU vendor.
  • Vulkan/CPU are compatibility fallbacks, never reported as vendor-native acceleration.
  • No GPU driver is replaced automatically. Installed drivers are qualified and reported as prerequisites.

Scenario and exact commands

Recommended product-level run through the protected Windows Dev Config bootstrap:

$url = 'https://raw.githubusercontent.com/microsoft/WindowsDeveloperConfig/main/src/windows-dev-config/bootstrap.ps1'
& ([scriptblock]::Create((irm $url))) -Scenario local-ai

For this unsigned PR branch before the sign cycle:

# First apply the documented temporary CurrentUser Bypass policy, then restore it.
$prHead = gh pr view 104 --repo microsoft/WindowsDeveloperConfig `
  --json headRefOid --jq .headRefOid
$prUrl = "https://raw.githubusercontent.com/microsoft/WindowsDeveloperConfig/$prHead/src/windows-dev-config/bootstrap.ps1"
& ([scriptblock]::Create((irm $prUrl))) `
  -Ref $prHead -Scenario local-ai -AllowUnsigned

The dispatcher downloads, verifies, protects, and copies the complete Workloads
tree plus shared Windows Dev Config helper steps. Microsoft-signed PowerShell
authenticates a content-hash manifest for every catalog/Python/C++/CUDA/config
input, and bootstrap reverifies the protected copy before launch. It elevates
the apply run and launches only the local-AI scenario—not the full workstation
setup. Unsigned branch tests use %ProgramData%\CalmOS-Development; the
production signed %ProgramData%\CalmOS payload remains isolated.

Repository-level equivalent:

.\Workloads\local-ai\install.ps1

Expected sentinels include:

PYTORCH_SMOKE=... "tensor_operation_verified": true, "model_forward_verified": true ...
PYTORCH_READY: backend=<CPU|CUDA|ROCm|XPU>, ...
LOCAL_AI_SCENARIO_READY: backend=Auto, runtime=None, ...

Select one optional runtime when needed:

.\Workloads\local-ai\install.ps1 -Runtime LlamaCpp
.\Workloads\local-ai\install.ps1 -Runtime Ollama
.\Workloads\local-ai\install.ps1 -Runtime Foundry

Product-level plan/runtime examples:

& ([scriptblock]::Create((irm $url))) `
  -Scenario local-ai -PlanOnly `
  -ReportRoot "$env:TEMP\local-ai-plan"

& ([scriptblock]::Create((irm $url))) `
  -Scenario local-ai -AiBackend CUDA -RequireTriton `
  -AiRuntime LlamaCpp

Dependency graph and transitive acquisition

local-ai
  -> hardware inventory
  -> pytorch(auto)
       -> NVIDIA: CUDA torch runtime
          -> when Triton selected: triton-windows + architecture-native MSVC + standalone CUDA Toolkit
       -> AMD: exact ROCm device runtime tuple inside PyTorch venv (not native ROCm SDK/hipcc)
       -> Intel: XPU tuple + triton-xpu (not full oneAPI)
       -> CPU: CPU tuple only
  -> optional one runtime:
       -> llama.cpp: backend assets + quick GGUF (NVIDIA assets include cudart)
       -> Ollama: source-managed backend/model, actual allocation reported
       -> Foundry: source-managed EP/model, actual EP/fallback reported

Drivers are prerequisites and are never replaced.

Entry point What it installs transitively / deliberately does not install
local-ai Inventory + PyTorch Auto with the exact vendor behavior above.
local-ai -Runtime LlamaCpp PyTorch stack + one backend-specific llama.cpp runtime + quick validation model; no independent full CUDA Toolkit requirement for llama assets.
local-ai -Runtime Ollama PyTorch stack + Ollama. x64 uses the installed WinGet application; ARM64 uses the official native archive converted into a managed per-user application with PATH/startup/upgrade/uninstall metadata. Actual backend/allocation is reported.
local-ai -Runtime Foundry PyTorch stack + source-managed Foundry EP/model; actual EP/fallback reported.
pytorch Contained backend runtime; CUDA/MSVC only when supported Triton JIT needs them; AMD does not install native hipcc; XPU does not install full oneAPI.
cuda Native CUDA Toolkit + architecture-native MSVC + kernel; no PyTorch/model runtime.
rocm Native ROCm SDK/hipcc + HIP kernel; no PyTorch/model runtime.
intel-ai OpenVINO CPU/GPU/NPU and optional full oneAPI/SYCL; no PyTorch XPU.
llama.cpp One backend runtime + quick GGUF; no PyTorch or unrelated runtime.
ollama Ollama + verified quick model. x64 is a normal registered application; ARM64 installs under %LOCALAPPDATA%\Programs\Ollama, never emulates the x64 setup EXE or selects a portable WinGet payload, and preserves models by default on uninstall.
foundry Foundry + catalog model; EP selection/acquisition is source-managed.

Standalone native-development flows remain available:

.\Workloads\cuda\install.ps1
.\Workloads\rocm\install.ps1
.\Workloads\intel-ai\install.ps1 -Device GPU -Profile Full

Quick local-model validation uses small Qwen models and real inference. An optional coding demonstration stays out of the default install:

.\Workloads\llama.cpp\coding-demo.ps1 `
  -ReportPath "$env:TEMP\llama-coding-demo.json"

It downloads the pinned Apache-2.0 Qwen2.5-Coder-1.5B-Instruct Q4_K_M GGUF (1,117,320,768 bytes, immutable revision and SHA-256), generates a Python group_anagrams implementation, and emits CODING_DEMO_READY.

Acquisition behavior

Component Windows acquisition What is installed automatically
PyTorch Official PyTorch/NVIDIA/AMD indexes or pinned wheel, selected by architecture/vendor/device Contained venv and required backend runtime tuple only
Triton CUDA Exact triton-windows==3.8.0.post28 community-stable build MSVC/CUDA toolchain only when the selected Triton JIT path needs it
Triton XPU Official PyTorch XPU integrated triton-xpu==3.8.0 Contained XPU tuple; no full oneAPI
CUDA WinGet x64; qualified direct ARM64 interim Native nvcc/MSVC kernel-development flow
ROCm/HIP AMD stable feed x64, exact supported gfx tuple Native hipcc/kernel-development flow
Intel AI Official OpenVINO Python tuple; optional WinGet oneAPI Requested OpenVINO CPU/GPU/NPU and/or native SYCL GPU tooling
llama.cpp Official backend-specific rolling assets from one release Selected CUDA/ROCm/SYCL/OpenVINO/Vulkan/OpenCL/CPU runtime
Ollama WinGet x64; latest official native ARM64 archive with GitHub digest x64 registered application; ARM64 atomic managed install, user PATH, HKCU startup, install manifest, model-preserving uninstall, and reported inference allocation
Foundry Qualified architecture-native WinGet preview Source-managed WinML/WebGPU/CUDA/CPU provider, reported after inference

Stable-channel status

Component / tuple Current acquisition Maturity/support status Stable channel available? Why not selected / promotion trigger Qualification required Tracking
CUDA x64 WinGet Nvidia.CUDA Stable Yes Selected Native kernel Current
CUDA ARM64 Pinned NVIDIA 13.4.0 prerelease installer Qualified interim developer preview Candidate: official signed direct 13.4.1, SHA-256 39af79e5e136c4e0de03bba816bda60fd7b70aad033e37ecaacf9f2e2c982442, 3,711,598,920 bytes Candidate has not passed the N1X workload suite nvcc compile/kernel plus PyTorch Triton JIT Tracked
PyTorch CUDA x64 Official cu126/cu130 indexes Stable Yes Selected by driver/capability Tensor/model + Triton Current
PyTorch CUDA ARM64 Pinned NVIDIA 2.15.0.dev...+cu134 wheel Qualified interim nightly Candidate: NVIDIA stable out-of-tree nvtorch_oot torch 2.14.0 + torchvision 0.29.0 + torchaudio 2.11.0 Exact trio is published but not N1X-qualified Trio imports, CUDA tensor/model, idempotence, Triton vector-add Tracked
Triton Windows CUDA PyPI triton-windows==3.8.0.post28 Community-stable, not upstream-official No upstream Windows package Retain exact qualified build Vector-add on each CUDA tuple Current / monitor
Triton XPU Official PyTorch XPU index Stable integrated Yes Selected Cold torch.compile on Intel GPU Current
Foundry WinGet Microsoft.FoundryLocal 0.10.3 Qualified preview Candidate: official non-prerelease v2.0.1; Python package metadata remains alpha v2 changes the CLI/SDK contract and is not yet target-qualified x64/ARM64 install, EP registration, real inference, cached rerun Tracked
llama.cpp Official rolling backend assets; Qualcomm pinned to policy-approved b10917 Rolling No backend-complete stable Windows channel WinGet exposes only x64 Vulkan Backend/device, actual layer offload, inference, policy acceptance Tracked
Ollama x64 WinGet Ollama.Ollama Stable installed application Yes Selected API/version/model/backend evidence Current
Ollama ARM64 Latest official ollama-windows-arm64.zip, converted into a managed Dev Config application Official native ARM64 archive; managed lifecycle supplied by Dev Config No official ARM64 setup EXE/current managed WinGet payload Promote when an upstream managed ARM64 installer/package passes native acceptance Native PE, digest, API/model/allocation, startup/PATH, upgrade/uninstall Current direct path / tracked promotion
ROCm x64 AMD stable feed Stable Yes Selected; WinGet ID unconfirmed Native HIP kernel Current
Intel OpenVINO / oneAPI x64 Official PyPI / WinGet Stable Yes Selected Requested-device inference / SYCL kernel Current

Artifact existence is not sufficient for promotion. Candidate metadata, hashes, triggers, and tracking state are centralized in src/Workloads/_common/ai-catalog.psd1, so promotion remains a resolver/data change after real workload qualification.

Hardware and backend coverage

  • NVIDIA: CUDA toolkit, PyTorch CUDA/Triton, llama.cpp CUDA.
  • AMD: ROCm/HIP, contained PyTorch ROCm tuple, llama.cpp ROCm on AMD's published Windows GPU/gfx matrix. Native Windows AMD Triton is unavailable. ROCm is GPU/HIP, not Ryzen AI NPU.
  • Intel: OpenVINO CPU/GPU/NPU; optional oneAPI/SYCL GPU; PyTorch XPU/Triton XPU; llama.cpp SYCL/OpenVINO. XPU and SYCL are GPU paths, not NPU.
  • Qualcomm/Adreno ARM64: llama.cpp OpenCL; Foundry is the vendor-neutral source-managed path. PyTorch has no native Qualcomm Windows accelerator backend.
  • Foundry/Ollama: source-managed acceleration. Setup reports the actual provider/device/allocation and accepts truthful CPU fallback where the runtime selects it.

Same-vendor adapter targeting uses CUDA/ROCm/explicit PyTorch -DeviceIndex, llama.cpp -Device, OpenVINO -OpenVinoDeviceId, or oneAPI -SyclDeviceSelector.

Real hardware evidence

NVIDIA RTX Spark N1X, Windows ARM64

  • CUDA 13.4 / nvcc V13.4.46, driver 616.62, CC12.1: compiled and executed native kernel.
  • PyTorch 2.15.0.dev20260904+cu134, NumPy 2.5.2: CUDA tensor and minimal neural model forward on the N1X.
  • triton-windows 3.8.0.post28: JIT vector-add kernel.
  • llama.cpp b10883 CUDA: physical gpu_info, backends=CUDA, 29/29 actual layers offloaded, real Qwen inference.
  • Optional Qwen2.5-Coder-1.5B demo: generated the requested Python implementation at 101.7 generation tokens/s.
  • Foundry qwen3-0.6b: real inference, truthful CPUExecutionProvider fallback on the tested catalog/runtime.
  • Ollama 0.34.4: official ollama-windows-arm64.zip installed under %LOCALAPPDATA%\Programs\Ollama; native ARM64 PE, release digest b49aa49306da9bae5b7498f320f76bd2a13739c3105bcb6e115a26f0dfd10d67, install manifest, PATH/HKCU startup, persistent 127.0.0.1:11434 endpoint, verified qwen3:0.6b digest, real inference, and /api/ps 100% GPU. Idempotent rerun returned already-current; model-preserving uninstall and reinstall passed.

Qualcomm Adreno X1-85, Windows ARM64

  • llama.cpp OpenCL: result.ready=true, 29/29 layers offloaded, real Qwen inference.
  • Foundry: result.ready=true, WebGPUExecutionProvider, selected GPU, no fallback.
  • Qualcomm runtime pinned to physically validated/policy-approved b10917 because managed Defender ASR blocked the newer unsigned b10919 executable; Defender was not bypassed.

Partner hardware remains pending for NVIDIA x64, AMD x64, and Intel x64 GPU/NPU.

Signed production and AllSigned testing

  • Production bootstrap explicitly verifies Microsoft signatures, installs to a protected directory, unblocks the verified files, and requests process-scoped RemoteSigned, matching current main.
  • src/tests/ai-common/all-signed.ps1 proves unsigned src/ entry points are rejected under AllSigned and valid Microsoft-signed release workloads load in Windows PowerShell 5.1 and PowerShell 7 with per-process publisher consent.
  • After the signing pipeline publishes top-level AI copies, the same test automatically executes every signed AI -PlanOnly entry point in both hosts.
  • Contributor/partner runs from unsigned src/ explicitly follow the documented temporary CurrentUser Bypass procedure and restore the prior policy afterward.

Validation

  • 1,475 hardware-independent assertions passed independently in PowerShell 7 and Windows PowerShell 5.1.
  • AllSigned contract passed in both PowerShell hosts.
  • All seven workload plans plus the scenario hardware/PyTorch plans produced nine schema-valid JSON reports.
  • The published bootstrap.ps1 -Scenario local-ai -AllowUnsigned -PlanOnly path completed through UAC, protected copy, content verification, and child routing with no blockers (using the documented temporary CurrentUser Bypass in both PowerShell hosts, then restoring it).
  • All source PowerShell parsed; Python smokes compiled; manifest, workflow YAML, and report schema parsed.
  • Command Palette ARM64 Debug build passed (warnings only; NuGet vulnerability feed and x64 IID optimizer were unavailable on the ARM64 host).
  • git diff --check passed.

Remaining gaps

  • Live partner runs: NVIDIA x64; AMD x64 ROCm/PyTorch/llama; Intel x64 OpenVINO CPU/GPU/NPU, oneAPI/SYCL, PyTorch XPU/Triton XPU, and llama SYCL/OpenVINO.
  • Upstream-unavailable: native Windows ARM64 ROCm/XPU, Qualcomm PyTorch acceleration, native Windows AMD Triton, generic ARM GPU toolkit, AMD Ryzen AI NPU flow, and unlisted Windows GPU stacks without authoritative artifacts.
  • Tracked promotions requiring qualification: CUDA ARM64 13.4.1, NVIDIA stable ARM64 PyTorch trio, and Foundry v2.0.1.

Michael Von Hippel and others added 18 commits September 29, 2026 10:24
Add independent CUDA, Foundry Local, PyTorch, llama.cpp, and Ollama flows with architecture-aware installation, model-free smoke tests, shared decision helpers, unit coverage, and documentation.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Support CUDA and GPU-backed PyTorch on RTX Spark ARM64, resolve current llama.cpp ARM64 assets, and make every AI flow run an end-to-end kernel or model inference by default.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Retry locked installer cleanup without masking results, harden ARM64 MSVC discovery, and update llama.cpp b10867 inference arguments and diagnostics.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Use the verified Build Tools bootstrapper in modify mode, wait for installer completion, and require the architecture-native compiler before CUDA or Triton setup continues.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Put the Visual Studio Installer directory on the child command PATH so VsDevCmd can resolve its bundled vswhere.exe under hardened executable lookup policies.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Validate exact installed package versions before pip work, cache the pinned ARM64 wheel after one verified download, and keep tensor and Triton readiness probes on every run.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Replace AI DSC acquisition with PR microsoft#93-style PowerShell setup, add AMD ROCm and Intel AI flows, centralize provider promotion metadata, and emit portable hardware acceptance reports.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Distinguish absent, outdated, and current packages; use exact upgrade operations with module-to-CLI fallback; and make package evidence schema-tolerant.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Use explicit native command results instead of inherited LASTEXITCODE and make Ollama managed-process cleanup tolerant of empty and already-exited processes.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Parse llama benchmark JSON after diagnostic prefixes, retain raw backend diagnostics, normalize Foundry cache paths, and force UTF-8 standalone capture.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Add exhaustive supported-cell resolution, vendor-native llama.cpp backends, self-contained PyTorch ROCm/XPU paths, deterministic adapter selection, and strict hardware evidence reporting.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Add a self-contained plan/apply/report workflow, complete vendor assignments, evidence return criteria, and final N1X Ollama acceptance.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Bound native executions, warm up OpenCL before inference, correct provider and JSON parsing, and pin the policy-approved Qualcomm llama.cpp build. Update partner commands and regression coverage.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 0c9c6f4f-4ef5-4ca6-acda-43c5632d325f
Verify unsigned source is blocked, Microsoft-signed release scripts load in Windows PowerShell and PowerShell 7, and signed AI flows are exercised automatically after the sign cycle.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Lead with hardware-selected PyTorch and optional model runtimes, add a pinned coding demo and neural forward smoke, track stable-channel candidates behind real workload qualification, and validate the AllSigned production contract.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Route local-ai through the verified Windows Dev Config bootstrap, authenticate non-PowerShell workload content with a signed hash manifest, propagate scenario blockers, and document every transitive acquisition.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Restore the protected local AI dispatcher on the latest bootstrap architecture and convert the official ARM64 archive into an atomic per-user application with PATH, startup, upgrade, uninstall, and live readiness evidence.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Update the executable capability matrix and signed content hash for the managed ARM64 startup/install tuple.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@Kixantrix
Michael Von Hippel (Kixantrix) force-pushed the mihippel-microsoft-windows-ai-setup-workloads branch from 8ea8bb1 to ca06e37 Compare September 29, 2026 17:52
Michael Von Hippel and others added 5 commits September 29, 2026 10:55
Import Microsoft.PowerShell.Utility explicitly so the protected Windows PowerShell scenario does not depend on first-use module autoloading.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Capture verified child bootstrap output across UAC so scenario and workstation handoff failures remain actionable without changing the signed-file relaunch model.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Reject unsigned and hash-mismatched release scripts, require the Microsoft signer, and use actual AllSigned execution to handle hosted-runner certificate-chain differences.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Use actual host execution as the signature contract and validate the Microsoft signer whenever the hosted Authenticode API exposes its certificate.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep development dispatcher files out of the signed production tree and normalize signed release probes to the repository's CRLF release contract before strict AllSigned validation.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant