memlnaut-nisps/ALIGNMENT.md
monkey-w1n5t0n 68d4cc4017 build(firmware): migrate to PlatformIO and vendor memllib (plan §5)
One cut, no dual path. Closes ALIGNMENT defect 3 ("Arduino-CLI build
machinery is actively hostile") and vision bullet 4.

platformio.ini carries 16 [env:], one per variant, each passing
-DMEMLNAUT_MODE_TYPE; selftest passes -DNISPS_SELFTEST=1 instead. The env list
IS the registry now — the .ino comment-registry and the NISPS_ST_* token-paste
table are deleted rather than migrated. L12 noted that table was already
silently missing the currently-shipped SLPWorkshop variant, which is the whole
argument against having a second list.

Also deleted: the Python/sed machinery that rewrote the COMMITTED .ino on every
build, the sketch symlink forest, the global TFT_eSPI User_Setup.h mutation
(now -D flags — TFT_eSPI's own documented PlatformIO recipe), the UF2
boot-mount detection stack (upload_protocol=picotool talks to the bootloader
directly), and build-firmware-arch.sh entirely. Scripts 683 -> 435 lines.

memllib is vendored at lib/memllib/ from upstream e291192; no submodules
remain. VENDORED.md records provenance and the re-sync procedure.

S9: a firmware-build CI job compiles three representative envs against a cached
toolchain and reports per-variant flash/RAM. Firmware is in an automated gate
for the FIRST time. The old ci.yml comment justified excluding it as "low
verification value" — an assessment that did not survive contact, since the
SelfTest variant sat broken for an unknown period calling a DisplayDriver
method that did not exist at the pinned memllib commit, and nothing noticed
because nothing built it.

Verified: all 16 envs build from an empty cache, each within ~520 bytes of the
arduino-cli binary it replaces, flash and RAM. Measured as .text+.rodata /
.data+.bss+vector+uninitialized — NOT PlatformIO's console line, which
double-counts .data on this board. This does not prove the hardware boots; no
flash+smoke test was possible and that stays an operator chokepoint.

  slpworkshop 248232/145028   pafsynth 256880/149716   selftest 216228/17960
  (all 16 in the CI log format; none exceeds 2% of a 16 MB flash)

Two traps recorded so nobody rediscovers them: vendoring memllib's subdirs
without a src/ wrapper makes PlatformIO's library builder silently compile
NOTHING while still linking; and project build_flags land BEFORE the
framework's own -std=gnu++17 -Os, so build_unflags is required.

CORRECTION carried in this commit: the firmware sizes in c19d846's message and
the first version of the memllib recon doc were wrong — SLPWorkshop 145348,
PAFSynth 145300, SelfTest 141840. They came from building variants in sequence
through a SHARED incremental arduino-cli build directory, which reused stale
objects and under-reported by ~75 KB. Clean-cache rebuilds of the identical
commit give 216736/18492 for SelfTest. The real cost of the memllib upstream
bump is +216 bytes flash, not +316. Never measure firmware size through a
reused build dir.

HISTORY NOTE: this commit and the docs commit before it were rebuilt (force-push,
2026-07-21) so that each contains only what its message describes. The first
versions had the firmware deletions stranded in the docs commit by a shared-index
race between concurrent agents; content is byte-identical to the originals.

Gates: run-all-tests.sh ALL GREEN (nisps/ untouched by this change beyond
include paths); 16/16 pio envs build.
2026-07-21 20:17:58 +02:00

13 KiB
Raw Blame History

ALIGNMENT

Opinionated diagnosis of how well the codebase serves its mission, ranked by impact. Dated entries; remove when resolved rather than checking off. Pruned every few weeks — a stale diagnosis is worse than none.

Mission

A research platform for interactive ML control of audio. We're building it to figure out what works and what doesn't — different ergonomics and ergodynamics of parameter sets, modes, ML architectures, audio engines, UI, and UX. Therefore: keep most/all parameters tweakable, ML/engine/UI/UX should each be configurable on their own axis, and the codebase has to enable/assist agentic AI coding patterns (confident changes, verifiable without hardware).

Target vision (operator, 2026-07-20): (1) one C++20 NISPS core serving RP2350 firmware and the browser, performance-sensitive on the MCU; (2) firmware modes runnable as modes in Manifold; (3) Manifold defaults to curated presets/modes, with the maximalist surface behind an "advanced" dev mode used to author them; (4) PlatformIO for hardware, no more .ino; (5) Manifold doubles as interface/editor for the hardware MEMLNaut (settings, presets, training, examples, visualisation).

The clean-slate rewrite (2026-04-29) consolidated everything into one C++20 codebase compiling to firmware AND WASM. Since 2026-07-13 (P1) the sole browser app is the React Manifold. JSON schemas remain the firmware↔browser parameter contract. A full-repo audit (2026-07-21, docs/specs/recon/simplification-audit-2026-07.md) grounds the entries below; mitigations are phased in docs/specs/plans/simplification-plan.md.

Top defects (ranked by mission impact)

1. The mode layer is not shared: WASM re-orchestrates modes by hand (2026-07-21)

What. nisps/modes/ — the CRTP layer binding ML config, engine, voice-space and I/O — compiles only into firmware. nisps/wasm/bindings.cpp includes engines and ML primitives but zero mode headers, and Manifold re-assembles mode behaviour (jolt stepping, OU, routing) in TS. "Firmware and WASM share the same modes" is true only at the engine level; every ModeBase behaviour must be mirrored browser-side by hand.

Why it blocks the mission. Vision bullet 2 is precisely this. Until the control-tick orchestration exists once in C++, every new mode behaviour is a dual implementation with drift risk.

Rough cost. Spec first, then ~a week: storage-policy the ModeBase orchestration the way P2 did MLPCore (verified shape in plan §6.5a — not binding monolithic mode objects, which would contradict the locked two-instance RT architecture). Related honesty gap: Manifold currently catalogues 4 modes that structurally cannot run in the browser (no mic input, event-only engines) — plan §6.5b (absorbs the old C15/mic-input defect; C15 itself lives on archive/playground-solidjs).

2. No curated/advanced split and no in-UI mode picker — the UI fights vision 3 (2026-07-21)

What. Manifold is 100% dev-maximalist: five drawers of everything, no preset data model to author against, and mode switching exists only via the debug hook — there is no instrument picker in the UI at all (the plumbing, ctx.modes/setModeId, already exists unused). A stratum of decorative controls (training-param sliders, master volume, bpm, A/B, snapshots, fabricated gradient health) renders real-looking UI that drives nothing.

Why it blocks the mission. The default experience is supposed to be curated presets; the advanced surface is the authoring tool. Neither exists, and the decorative stratum actively misleads research use.

Rough cost. Product-model decision first (plan §7.6), then incremental: picker is days; the curated-preset model seeds from backends/presets.ts + schemas; disclosure via per-drawer depth levels. Deleting the decorative stratum is part of the Phase-1 sweep.

3. Manifold-as-hardware-editor is a facade (2026-07-21)

What. Vision bullet 5 exists as a 237-line Web Serial shell: sound connect lifecycle, zero protocol (saveModel/restoreModel/getSettings are literal stubs), and firmware has no serial command surface or on-device persistence to talk to.

Why it blocks the mission. The hardware research loop (train on device, inspect/curate in browser) is closed only by this bridge.

Rough cost. Week+, spec-first (plan §6.5d). The right discipline already exists in-repo: useq-celium's C-header wire truth + TS mirror + parity test; settings payloads should derive from schema codegen.

4. Dead mass and registry sprawl across every layer (2026-07-21)

What. Phase 1 landed 2026-07-21 and removed the bulk of this: the dead focus/altitude UI system, the decorative control stratum, 12 dead WASM API entries across the 5-file registration chain, the vendored daisysp tree, retired-playground artifacts and root planning relics, 5 unused primitives, the duplicate backend editor and catalogue, voice_space.hpp, fixed_buffer.hpp, the dead perf-macro regime, and the OSC bridge twin. What remains is the registry half: mode identity spread across ~6 hand-maintained registries with demonstrated drift, MLP dims typed twice, per-mode schema blocks hand-written in C++, and assorted stale specs presenting a deleted world as present tense.

Why it blocks the mission. The registries are dual-truth bugs waiting to fire (one already did: the selftest table). Stale specs are agent-confusing surface area.

Rough cost. Plan phase 3 (~23 days, codegen takes ownership) plus the docs disposition pass (§8). The behaviour bugs found en route (dataset-cap divergence, VCV 2-D input truncation, VCV audio-thread race and JSON) were fixed in phase 2 on 2026-07-21.

5. No performance measurement despite a performance-defined mission (2026-07-21)

What. The "super performance-sensitive" constraint is enforced only by static discipline (the no-heap lint — false negatives closed in Phase 2 — and section attrs). Half-closed 2026-07-21: the Phase 4 firmware CI job now reports per-variant flash/RAM on every push, so size regressions are at least visible. Still missing: any measure of time. No benchmark, no CPU-load assertion, no blocks-per-second number on either target — nothing would catch an engine getting 3x slower.

Rough cost. ~Half a day now: a host-side blocks-per-second benchmark for engine_process_block, native + WASM (plan §6.5f).

6. Training-health telemetry: decided, not yet built (2026-07-21)

What. Four fragments of one feature. Fragment 3 (decorative gradient-health UI) was deleted in Phase 1. The other three stand: a 16 KB loss-history buffer in every firmware MLP that nothing reads (nisps/ml/mlp.hpp); a WASM worker faking a 1-element loss history (manifold/src/engine/wasm-worker.ts:310, new Float32Array([loss])); and a real layer_stats / nisps_ml_get_layer_stats API plumbed end-to-end and consumed by nobody.

Decisions are now complete (operator, §7.3 + L25): telemetry becomes real, browser-only, behind a feature flag; the fakes go; and the firmware buffer stays — it is the on-device record the hardware editor (defect 3) will want, and it costs flash we demonstrably have (16% RAM on the largest variant). So the remaining work is one coherent job, not a judgement call: plumb loss_history through the C API (Drawers.tsx:263 already marks the gap), replace the worker's fake with it, and put the display plus get_layer_stats behind the advanced-mode flag.

Why it blocks the mission. "Is the network learning?" is a core research affordance, and today it looks answered while being fabricated — worse than absent.

Rough cost. ~A day, spec-light: plan §6.5e, no longer gated on anything.

7. RMSProp still deferred from nisps/ml/ (2026-04-29; reaffirmed 2026-07-21)

What. training.hpp ships SGD only; the legacy firmware used RMSProp for TrainBatch. Optimizer choice is a research axis. Not blocking current fits; will matter for harder loss landscapes. Port target: upstream MusicallyEmbodiedML memlp (the in-repo src/memlp copy is deleted; use the GitHub remote or archive branch).

Rough cost. A day, plus batch-convergence tests.

Open mission questions

Q1: Per-mode MLP architectures or one shared shape? (2026-04-29)

Schemas declare per-mode dims and since P5.3 both targets honour them. Is the mission served by maintaining per-mode shapes (research diversity) or collapsing to one (simpler ops)? Note the audit found all 9 mode schemas share copy-pasted ML defaults and 20 params are anonymous placeholders — the per-mode diversity is currently nominal (plan L40).

Q2: Engine event taxonomy (2026-04-29)

ControlEvent is a flat enum consumed by the two sequencer modes. Revisit when a third event-emitting mode lands.

Q3: Should Manifold stay desktop-first? (2026-04-29)

Legacy a-immersive was mobile-first; Manifold is desktop-first. Defer until user data exists.

Q4: Who owns memllib? — DECIDED, half-executed (2026-07-21)

Operator decision: vendor, self-contained in this repo. The inventory (docs/specs/recon/memllib-usage-inventory.md) settled the shape: there is no small load-bearing subset — it is all of memllib bar examples/ (~1.8 MB, 24/24 compiled TUs link). The fork is dissolved: its three commits touch only examples/, which the firmware never compiles and whose content already lives in nisps/ml/{jolt,ou_noise,feedback,geo_push}.hpp, so the submodule now points at upstream and is pinned to current main. Verified by building: it brings the l r input swap hardware fix plus the NavigateToView the SelfTest variant was already written against, and costs +216 bytes of a 16 MB flash. Remaining: the vendoring copy itself, which lands with the PlatformIO cut (plan §5). Delete this entry when it does.

Q5: Legacy feedback modes — delete or keep for A/B? (2026-07-21)

RandomiseOutputs/RandomiseMlp/Diffuse/on_drag have no product consumer, but docs/adr/rl-feedback-design.md explicitly kept Diffuse for A/B comparison. Deleting reverses a recorded decision — operator call (plan §7.1).

Deferred / accepted debt

  • EOC effects chain, ShapeSeq sequencer, modular engine (Phase E) — legacy features consciously out of the v1 rewrite; revisit only if a mode wants them.
  • Inputs multi-source composition (2026-06-28, reaffirmed 2026-07-21) — mix-and-match pad+gamepad+MIDI is a recorded, unreversed decision; the UI currently enforces exclusive single-source and the composition machinery sits dormant by design. Schedule or keep dormant — but the inputs-spec must stop presenting composition as current behaviour (plan §8).
  • Schema content is partially placeholder (2026-07-21) — 20 anonymous "Param NN" slots across paf_synth/channel_strip/xiasri and copy-pasted ML defaults across all 9 modes. Name them during the first curated-preset pass per mode (plan §6.5c), or shrink output_size where the engine allows.
  • Geometric-dislike deliberate divergences (2026-07-14, one-core P3): (1) the degenerate-branch RNG draws from the controller's deterministic nisps::Rng, not libc rand() — native==WASM parity holds; (2) upstream's async shuffled two-LR optimise() is collapsed into one synchronous dislike_geometric() training only the pressed negative's target — behavioural, not bitwise, parity with firmware upstream, by design; (3) RandomiseMlp uses draw_weights(spread) rather than the old asymmetric ranges. All intentional.
  • Manifold dock splits state/muted/armed (2026-06-28) — deliberate divergence from the deployed conflated frozenmuted model (dock-spec §3.3). muted-downstream and the soloMode gradient-mask variants remain UI-only; the C API exposes set_focus but no per-mode gradient masking yet. (The audit found soloMode behaviourally inert in the controller — plan L20 trims it until train_masked exists.)

Recently resolved (delete after a few weeks)

  • 2026-07-21: Arduino-CLI build machinery (old defect 3) is gone. Phase 4 replaced it with a PlatformIO project: one [env:] per variant is now the only variant registry, the .ino-mutating Python/sed machinery and the NISPS_ST_* token-paste table and the sketch symlink forest and the global TFT_eSPI mutation are all deleted, memllib is vendored (no submodule), and firmware finally entered CI — three representative envs per run, which is what would have caught the SelfTest variant sitting broken. All 16 envs build; sizes match arduino-cli within ~520 bytes.

  • 2026-07-21: Full-repo simplification audit landed (recon + plan + this rewrite). Superseded entries removed: "browser-only engines incomplete" (→ defect 2/plan 5b), "loss curve not plumbed" (→ defect 7), "NISPS_AUDIO_FUNC misshapen" (→ plan Phase 1, S21/L13), stale "VCV not currently maintained" note (vcv/ is active and consumes nisps/ directly post-P6).

  • 2026-07-18: Browser curve maths unified onto the canonical nisps/core/math.hpp catalog at P4; four silently-divergent TS curves re-baselined.

  • 2026-07-14: WASM MLP fixed-architecture defect resolved by P2 (MLPCore<Storage>; browser runtime-shaped, firmware zero-heap fixed).