refactor(ml)!: P2.1 storage-policy split — MLPCore<Storage>, fixed + dynamic models
Algorithms (forward, backprop/SGD, init, move_weights, diagnostics) now live
once in MLPCore<Storage> (nisps/ml/mlp.hpp). Storage models:
- FixedStorage (storage.hpp): template-sized std::array, zero heap. The
classic MLP<NIn,H1,H2,H3,NOut,...> is an alias preserving kInput/kHidden*/
kOutput/kNumLayers/weight_count() constexpr — firmware + bindings + modes
compile unchanged.
- DynamicStorage (dynamic_storage.hpp): runtime dims, ONE arena allocation
at construction, nothing per-call. #error under NISPS_TARGET_EMBEDDED
(new macro in core/perf.hpp); sole lint-cpp.sh heap-allowlist entry, plus
a lint check that fails if the #error guard disappears.
Verification:
- new ctest test_mlp_storage_parity: fixed↔dynamic BIT-identical across
init/draw/inference/train(FIFO)/move_weights(pin mask)/eval_loss/
layer_stats/set_weights/infer_batch/reset; invalid+moved-from inert
- golden ML vectors (pre-refactor constants) pass → bit-stable refactor
- native↔WASM parity PASS, max delta unchanged (2.4e-7)
- chokepoint B compile: PAFSynth .text 122324→122692 (+0.30%, ±1% budget);
RAM +416B (eval scratch)
- fix: firmware-common.sh used bare 'python' (absent here) → ${PYTHON:-python3}
Part of one-core-engine-refactor P2. nisps_ml_create ABI untouched (P2.2 is
an operator stop-point).
This commit is contained in:
parent
43ef6d5dae
commit
8a19e5b52c
11 changed files with 900 additions and 343 deletions
2
MAP.md
2
MAP.md
|
|
@ -6,7 +6,7 @@ MEMLNaut-NISPS — Neural Interactive Shaping of Parameter Spaces. One C++20 cod
|
|||
|
||||
### `nisps/` — platform-agnostic C++20 library (the only ML/DSP/engine code)
|
||||
- `nisps/core/` — `perf.hpp` (memory section attrs), `types.hpp`, `concepts.hpp` (`MLEngine`, `AudioEngine`, `Mode`), `fixed_buffer.hpp`, `ring_buffer.hpp` (SPSC lock-free, replaces pico/util/queue), `rng.hpp` (xoshiro256+ deterministic), `math.hpp` (fast_sigmoid, `Curve` enum + `apply_curve`).
|
||||
- `nisps/ml/` — MLP class template `MLP<NIn, NH1, NH2, NH3, NOut>`. Files: `mlp.hpp`, `activations.hpp`, `loss.hpp` (MSE, no double-scaling), `training.hpp` (SGD + grad clipping), `init.hpp` (spread-aware uniform↔Xavier), `rl.hpp` (`move_weights` with output pin mask + per-layer scaling + weight decay), `jolt.hpp` (`Jolt` — held continuous weight-morph over the flat weight buffer + post-release LR ramp; ported from upstream InterfaceRL), `ou_noise.hpp` (`OUNoise<N>` — Ornstein-Uhlenbeck exploration walk on the output vector; ported from upstream InterfaceRL), `feedback.hpp` (`FeedbackController<MLP_T>` — the 3-mode "Down Action" negative-feedback state machine: Avoid / RandomiseOutputs / RandomiseMlp; header-only, zero-heap, own deterministic RNG, exposed via `nisps_ml_feedback_*` C API), `stats.hpp`. Jolt + OU are inert by default and wired into `ModeBase`, so every mode exposes `jolt_press/jolt_release`, `jolt_lr_scale`, and `set_explore_intensity`.
|
||||
- `nisps/ml/` — the MLP core, written once against a storage policy (`mlp.hpp` `MLPCore<Storage>`): `storage.hpp` (`FixedStorage` — template-sized `std::array`, zero heap; `MLP<NIn,NH1,NH2,NH3,NOut>` alias preserves the classic compile-time surface) and `dynamic_storage.hpp` (`DynamicStorage` — runtime dims, single arena alloc at construction; `#error`s on RP2350 builds, sole lint heap-allowlist entry). Fixed↔dynamic bit-parity enforced by `tests/cpp/test_mlp_storage_parity.cpp`. Files: `mlp.hpp`, `activations.hpp`, `loss.hpp` (MSE, no double-scaling), `training.hpp` (SGD + grad clipping), `init.hpp` (spread-aware uniform↔Xavier), `rl.hpp` (`move_weights` with output pin mask + per-layer scaling + weight decay), `jolt.hpp` (`Jolt` — held continuous weight-morph over the flat weight buffer + post-release LR ramp; ported from upstream InterfaceRL), `ou_noise.hpp` (`OUNoise<N>` — Ornstein-Uhlenbeck exploration walk on the output vector; ported from upstream InterfaceRL), `feedback.hpp` (`FeedbackController<MLP_T>` — the 3-mode "Down Action" negative-feedback state machine: Avoid / RandomiseOutputs / RandomiseMlp; header-only, zero-heap, own deterministic RNG, exposed via `nisps_ml_feedback_*` C API), `stats.hpp`. Jolt + OU are inert by default and wired into `ModeBase`, so every mode exposes `jolt_press/jolt_release`, `jolt_lr_scale`, and `set_explore_intensity`.
|
||||
- `nisps/dsp/` — `biquad.hpp`, `delay.hpp`, `reverb.hpp`, `filter.hpp`, `env.hpp`, `osc.hpp`, `pitch_shift.hpp`, `dc_blocker.hpp`. Lean primitives extracted from maximilian; daisysp PitchShifter replaced with custom granular impl.
|
||||
- `nisps/engines/` — eight audio engines, each satisfying `AudioEngine`: `paf_synth.hpp`, `channel_strip.hpp`, `xiasri.hpp`, `verb_fx.hpp`, `memlcelium.hpp`, `breakor.hpp` (sequencer, NoOp audio), `elysiamorf.hpp` (sequencer, NoOp audio), `analysis.hpp` (input-side spectral features). Plus `base.hpp` (`NoOpEngine`, engine_id "thru").
|
||||
- `nisps/modes/` — platform-agnostic modes binding `{ML config, engine, voice space lambdas, abstract I/O channels}`. Files: `paf_synth.hpp`, `channel_strip.hpp`, `xiasri.hpp`, `verb_fx.hpp`, `memlcelium.hpp`, `slp_workshop.hpp` (`SLPWorkshopMode` — the Synth Library Portland workshop build; reuses the MEMLCelium engine + MLP shape, foregrounds the Jolt + OU explore gestures), `breakor.hpp`, `elysiamorf.hpp`, `sound_analysis_midi.hpp`, `external_synth_midi.hpp` (`ExternalSynthMIDIMode<const MidiDevice&, NOut>` — joystick→MLP→MIDI CC for an external synth; compile-time device from `nisps/midi`; `consteval pick_cc_slots` curates which params fill the NOut slots; NoOpEngine, `kRouteOutputsToEngine=false`). `base.hpp` provides a CRTP scaffold eliminating the duplication that previously plagued firmware modes. `voice_space.hpp` holds engine-side voice space dispatch helpers. `generated/` contains codegen output (do not edit by hand).
|
||||
|
|
|
|||
|
|
@ -101,13 +101,15 @@ Each phase ends green on its test gate and is independently landable. File phase
|
|||
|
||||
### P2 — Storage-policy split: templated hardware, dynamic browser (the structural centre, ≈1 wk)
|
||||
|
||||
- Refactor `nisps/ml/` so algorithms (forward, backprop/SGD, init, `move_weights`, jolt, OU, feedback)
|
||||
are written once against a storage concept: `weights()`, `layer_sizes()`, `scratch()`. Two models:
|
||||
- `FixedStorage<NIn,H1,H2,H3,NOut>` — `std::array`, `NISPS_AUDIO_MEM`-able, zero heap. Firmware target;
|
||||
existing `MLP<...>` becomes an alias. **RP2350 performance contract untouched.**
|
||||
- `DynamicStorage` — sizes at construction, single arena allocation, no per-call allocation after
|
||||
construction. Compiled only for WASM/native-test/VCV targets (guarded so `lint-cpp.sh` still fails heap
|
||||
use in firmware paths).
|
||||
- ✅ (landed 2026-07-13) Refactor `nisps/ml/` so algorithms (forward, backprop/SGD, init, `move_weights`)
|
||||
are written once against a storage policy (`mlp.hpp` `MLPCore<Storage>`; jolt/OU/feedback already operate
|
||||
on the MLP surface and needed no change). Two models:
|
||||
- `FixedStorage<NIn,H1,H2,H3,NOut>` (`nisps/ml/storage.hpp`) — `std::array`, zero heap; `MLP<...>` is an
|
||||
alias preserving the full compile-time surface. **RP2350 contract verified: PAFSynth `.text`
|
||||
122324→122692 = +0.30% (±1% budget); ctest + golden vectors + WASM parity bit-stable.**
|
||||
- `DynamicStorage` (`nisps/ml/dynamic_storage.hpp`) — sizes at construction, single arena allocation,
|
||||
nothing per-call. `#error`s under `NISPS_TARGET_EMBEDDED`; sole `lint-cpp.sh` heap-allowlist entry, with
|
||||
a lint check that fails if the guard is ever removed.
|
||||
- `nisps_ml_create(input, output, hidden[])` honours its arguments. Reshape = new instance + warm-start
|
||||
copy of overlapping weights (the BUILD-PLAN warm-start idea, now runtime).
|
||||
- Manifold drops input clamping/phantom-channel handling; XIASRI/sound-analysis multi-input modes become
|
||||
|
|
|
|||
Binary file not shown.
|
|
@ -56,6 +56,7 @@ if(NOT EMSCRIPTEN)
|
|||
${NISPS_TEST_DIR}/test_mlp_training.cpp
|
||||
${NISPS_TEST_DIR}/test_mlp_loss.cpp
|
||||
${NISPS_TEST_DIR}/test_mlp_rl.cpp
|
||||
${NISPS_TEST_DIR}/test_mlp_storage_parity.cpp
|
||||
${NISPS_TEST_DIR}/test_mlp_jolt.cpp
|
||||
${NISPS_TEST_DIR}/test_mlp_ou_noise.cpp
|
||||
${NISPS_TEST_DIR}/test_mlp_feedback.cpp
|
||||
|
|
|
|||
|
|
@ -10,6 +10,15 @@
|
|||
|
||||
#pragma once
|
||||
|
||||
// NISPS_TARGET_EMBEDDED marks builds for the RP2350 hardware target. Code
|
||||
// that is allowed heap allocation at construction time on host/WASM targets
|
||||
// (e.g. nisps/ml/dynamic_storage.hpp) is compile-time excluded when this is
|
||||
// defined — the zero-heap firmware contract is enforced structurally, not
|
||||
// just by lint.
|
||||
#if defined(ARDUINO_ARCH_RP2040) || defined(ARDUINO_ARCH_RP2350)
|
||||
#define NISPS_TARGET_EMBEDDED 1
|
||||
#endif
|
||||
|
||||
#if defined(ARDUINO_ARCH_RP2040) || defined(ARDUINO_ARCH_RP2350)
|
||||
// Pico SDK provides __not_in_flash and __not_in_flash_func.
|
||||
// __not_in_flash takes a section name string; __not_in_flash_func wraps the
|
||||
|
|
|
|||
194
nisps/ml/dynamic_storage.hpp
Normal file
194
nisps/ml/dynamic_storage.hpp
Normal file
|
|
@ -0,0 +1,194 @@
|
|||
// nisps/ml/dynamic_storage.hpp — runtime-shaped storage policy for the MLP
|
||||
// core. WASM / native-test / VCV targets ONLY.
|
||||
//
|
||||
// Dimensions are chosen at construction; every buffer lives in ONE arena
|
||||
// allocated once in the constructor. There is NO allocation after
|
||||
// construction — the algorithm hot paths are as allocation-free as the
|
||||
// fixed model.
|
||||
//
|
||||
// This header is compile-time excluded from embedded (RP2350) builds: the
|
||||
// firmware's zero-heap contract is structural, not advisory. lint-cpp.sh
|
||||
// additionally allowlists exactly this file for its heap audit — heap use
|
||||
// anywhere else under nisps/ml/ still fails the lint.
|
||||
//
|
||||
// See nisps/ml/storage.hpp for the storage surface contract, and
|
||||
// docs/specs/plans/one-core-engine-refactor.md §P2 for the design.
|
||||
|
||||
#pragma once
|
||||
|
||||
#include "../core/perf.hpp"
|
||||
|
||||
#if defined(NISPS_TARGET_EMBEDDED)
|
||||
#error "nisps/ml/dynamic_storage.hpp must not be compiled for the RP2350 target (zero-heap contract)"
|
||||
#endif
|
||||
|
||||
#include <cstddef>
|
||||
#include <new>
|
||||
#include <span>
|
||||
|
||||
#include "storage.hpp" // kMlpNumLayers
|
||||
|
||||
namespace nisps::ml {
|
||||
|
||||
class DynamicStorage {
|
||||
public:
|
||||
static constexpr std::size_t kNumLayers = kMlpNumLayers;
|
||||
|
||||
// hidden must have exactly 3 entries (the 4-layer topology is fixed;
|
||||
// only the dimensions are runtime). All dims must be >= 1.
|
||||
DynamicStorage(std::size_t n_in,
|
||||
std::span<const std::size_t> hidden,
|
||||
std::size_t n_out,
|
||||
std::size_t max_examples = 128u,
|
||||
std::size_t max_iter_train = 4096u) noexcept {
|
||||
if (hidden.size() != 3u || n_in == 0u || n_out == 0u ||
|
||||
hidden[0] == 0u || hidden[1] == 0u || hidden[2] == 0u) {
|
||||
return; // stays invalid
|
||||
}
|
||||
dims_[0] = n_in;
|
||||
dims_[1] = hidden[0];
|
||||
dims_[2] = hidden[1];
|
||||
dims_[3] = hidden[2];
|
||||
dims_[4] = n_out;
|
||||
max_ex_ = max_examples;
|
||||
max_iter_ = max_iter_train;
|
||||
|
||||
std::size_t total = 0u;
|
||||
auto claim = [&total](std::size_t n) {
|
||||
const std::size_t off = total;
|
||||
total += n;
|
||||
return off;
|
||||
};
|
||||
for (std::size_t l = 0; l < kNumLayers; ++l) {
|
||||
off_w_[l] = claim(fan_in(l) * fan_out(l));
|
||||
off_b_[l] = claim(fan_out(l));
|
||||
off_pa_[l] = claim(fan_out(l));
|
||||
off_a_[l] = claim(fan_out(l));
|
||||
off_gw_[l] = claim(fan_in(l) * fan_out(l));
|
||||
off_gb_[l] = claim(fan_out(l));
|
||||
off_d_[l] = claim(fan_in(l));
|
||||
off_e_[l] = claim(fan_out(l));
|
||||
}
|
||||
off_input_ = claim(dims_[0]);
|
||||
off_output_ = claim(dims_[4]);
|
||||
off_dsf_ = claim(max_ex_ * dims_[0]);
|
||||
off_dsl_ = claim(max_ex_ * dims_[4]);
|
||||
off_flat_ = claim(weight_count());
|
||||
off_lh_ = claim(max_iter_);
|
||||
|
||||
// The single arena allocation. Value-init zeroes it, matching the
|
||||
// zero-initialised std::array members of FixedStorage.
|
||||
arena_ = new (std::nothrow) float[total]();
|
||||
total_ = (arena_ != nullptr) ? total : 0u;
|
||||
}
|
||||
|
||||
~DynamicStorage() { delete[] arena_; }
|
||||
|
||||
DynamicStorage(const DynamicStorage&) = delete;
|
||||
DynamicStorage& operator=(const DynamicStorage&) = delete;
|
||||
DynamicStorage(DynamicStorage&& o) noexcept { move_from_(o); }
|
||||
DynamicStorage& operator=(DynamicStorage&& o) noexcept {
|
||||
if (this != &o) {
|
||||
delete[] arena_;
|
||||
move_from_(o);
|
||||
}
|
||||
return *this;
|
||||
}
|
||||
|
||||
bool valid() const noexcept { return arena_ != nullptr; }
|
||||
|
||||
// ---- dims -----------------------------------------------------------
|
||||
std::size_t n_in() const noexcept { return dims_[0]; }
|
||||
std::size_t n_out() const noexcept { return dims_[4]; }
|
||||
std::size_t max_examples() const noexcept { return max_ex_; }
|
||||
std::size_t max_iter_train() const noexcept { return max_iter_; }
|
||||
|
||||
std::size_t fan_in(std::size_t l) const noexcept { return dims_[l]; }
|
||||
std::size_t fan_out(std::size_t l) const noexcept { return dims_[l + 1u]; }
|
||||
|
||||
template <std::size_t L> std::size_t fan_in_l() const noexcept { return dims_[L]; }
|
||||
template <std::size_t L> std::size_t fan_out_l() const noexcept { return dims_[L + 1u]; }
|
||||
|
||||
std::size_t weight_count() const noexcept {
|
||||
return dims_[0] * dims_[1] + dims_[1] * dims_[2] + dims_[2] * dims_[3] +
|
||||
dims_[3] * dims_[4] + dims_[1] + dims_[2] + dims_[3] + dims_[4];
|
||||
}
|
||||
|
||||
// ---- per-layer buffers ------------------------------------------------
|
||||
template <std::size_t L> std::span<float> weights_l() noexcept {
|
||||
return {arena_ + off_w_[L], fan_in_l<L>() * fan_out_l<L>()};
|
||||
}
|
||||
template <std::size_t L> std::span<const float> weights_l() const noexcept {
|
||||
return {arena_ + off_w_[L], fan_in_l<L>() * fan_out_l<L>()};
|
||||
}
|
||||
template <std::size_t L> std::span<float> biases_l() noexcept {
|
||||
return {arena_ + off_b_[L], fan_out_l<L>()};
|
||||
}
|
||||
template <std::size_t L> std::span<const float> biases_l() const noexcept {
|
||||
return {arena_ + off_b_[L], fan_out_l<L>()};
|
||||
}
|
||||
template <std::size_t L> std::span<float> pre_act_l() noexcept {
|
||||
return {arena_ + off_pa_[L], fan_out_l<L>()};
|
||||
}
|
||||
template <std::size_t L> std::span<float> act_l() noexcept {
|
||||
return {arena_ + off_a_[L], fan_out_l<L>()};
|
||||
}
|
||||
template <std::size_t L> std::span<const float> act_l() const noexcept {
|
||||
return {arena_ + off_a_[L], fan_out_l<L>()};
|
||||
}
|
||||
template <std::size_t L> std::span<float> grad_w_l() noexcept {
|
||||
return {arena_ + off_gw_[L], fan_in_l<L>() * fan_out_l<L>()};
|
||||
}
|
||||
template <std::size_t L> std::span<float> grad_b_l() noexcept {
|
||||
return {arena_ + off_gb_[L], fan_out_l<L>()};
|
||||
}
|
||||
template <std::size_t L> std::span<float> delta_l() noexcept {
|
||||
return {arena_ + off_d_[L], fan_in_l<L>()};
|
||||
}
|
||||
template <std::size_t L> std::span<float> eval_act_l() const noexcept {
|
||||
return {arena_ + off_e_[L], fan_out_l<L>()};
|
||||
}
|
||||
|
||||
// ---- global buffers ---------------------------------------------------
|
||||
std::span<float> input_buf() noexcept { return {arena_ + off_input_, dims_[0]}; }
|
||||
std::span<const float> input_buf() const noexcept { return {arena_ + off_input_, dims_[0]}; }
|
||||
std::span<float> output_buf() noexcept { return {arena_ + off_output_, dims_[4]}; }
|
||||
std::span<const float> output_buf() const noexcept { return {arena_ + off_output_, dims_[4]}; }
|
||||
std::span<float> ds_features() noexcept { return {arena_ + off_dsf_, max_ex_ * dims_[0]}; }
|
||||
std::span<const float> ds_features() const noexcept { return {arena_ + off_dsf_, max_ex_ * dims_[0]}; }
|
||||
std::span<float> ds_labels() noexcept { return {arena_ + off_dsl_, max_ex_ * dims_[4]}; }
|
||||
std::span<const float> ds_labels() const noexcept { return {arena_ + off_dsl_, max_ex_ * dims_[4]}; }
|
||||
std::span<float> flat_buf() noexcept { return {arena_ + off_flat_, weight_count()}; }
|
||||
std::span<float> loss_hist_buf() noexcept { return {arena_ + off_lh_, max_iter_}; }
|
||||
std::span<const float> loss_hist_buf() const noexcept { return {arena_ + off_lh_, max_iter_}; }
|
||||
|
||||
private:
|
||||
void move_from_(DynamicStorage& o) noexcept {
|
||||
for (std::size_t i = 0; i < 5u; ++i) dims_[i] = o.dims_[i];
|
||||
max_ex_ = o.max_ex_; max_iter_ = o.max_iter_;
|
||||
for (std::size_t l = 0; l < kNumLayers; ++l) {
|
||||
off_w_[l] = o.off_w_[l]; off_b_[l] = o.off_b_[l];
|
||||
off_pa_[l] = o.off_pa_[l]; off_a_[l] = o.off_a_[l];
|
||||
off_gw_[l] = o.off_gw_[l]; off_gb_[l] = o.off_gb_[l];
|
||||
off_d_[l] = o.off_d_[l]; off_e_[l] = o.off_e_[l];
|
||||
}
|
||||
off_input_ = o.off_input_; off_output_ = o.off_output_;
|
||||
off_dsf_ = o.off_dsf_; off_dsl_ = o.off_dsl_;
|
||||
off_flat_ = o.off_flat_; off_lh_ = o.off_lh_;
|
||||
arena_ = o.arena_; total_ = o.total_;
|
||||
o.arena_ = nullptr; o.total_ = 0u;
|
||||
}
|
||||
|
||||
std::size_t dims_[5] = {0u, 0u, 0u, 0u, 0u};
|
||||
std::size_t max_ex_ = 0u;
|
||||
std::size_t max_iter_ = 0u;
|
||||
std::size_t off_w_[kNumLayers]{}, off_b_[kNumLayers]{}, off_pa_[kNumLayers]{},
|
||||
off_a_[kNumLayers]{}, off_gw_[kNumLayers]{}, off_gb_[kNumLayers]{},
|
||||
off_d_[kNumLayers]{}, off_e_[kNumLayers]{};
|
||||
std::size_t off_input_ = 0u, off_output_ = 0u, off_dsf_ = 0u, off_dsl_ = 0u,
|
||||
off_flat_ = 0u, off_lh_ = 0u;
|
||||
float* arena_ = nullptr;
|
||||
std::size_t total_ = 0u;
|
||||
};
|
||||
|
||||
} // namespace nisps::ml
|
||||
641
nisps/ml/mlp.hpp
641
nisps/ml/mlp.hpp
|
|
@ -1,9 +1,8 @@
|
|||
// nisps/ml/mlp.hpp — fixed-architecture MLP, four layers (three hidden +
|
||||
// output). All buffers are template-sized std::array; zero heap allocation
|
||||
// in inference, training, and the dataset path.
|
||||
// nisps/ml/mlp.hpp — four-layer MLP (three hidden + output), written ONCE
|
||||
// against a storage policy (docs/specs/plans/one-core-engine-refactor.md P2).
|
||||
//
|
||||
// ARCHITECTURE
|
||||
// MLP<NIn, NHidden1, NHidden2, NHidden3, NOut, NMaxExamples = 128>
|
||||
// MLPCore<Storage>
|
||||
// ┌──────┐ Linear+Bias ┌────────┐ ReLU ┌────────┐ ReLU ┌────────┐ Sigmoid
|
||||
// │ NIn │ ─────────────▶ │ NH1 │ ──────▶ │ NH2 │ ──────▶ │ NH3 │ ──────▶ NOut
|
||||
// └──────┘ └────────┘ └────────┘ └────────┘
|
||||
|
|
@ -12,46 +11,35 @@
|
|||
// Layer 2 (NH2 → NH3) ReLU
|
||||
// Layer 3 (NH3 → NOut) Sigmoid
|
||||
//
|
||||
// We support exactly three hidden layers. The legacy firmware default is
|
||||
// [10, 10, 14], so the MVP signature directly matches `MLP<NIn, 10, 10, 14,
|
||||
// NOut>`. Variable layer count is deferred — see architecture.md.
|
||||
// The topology (4 layers, ReLU×3 + Sigmoid) is fixed; the DIMENSIONS come
|
||||
// from the storage policy:
|
||||
//
|
||||
// MEMORY MODEL
|
||||
// Per layer L_k with fan_in = N_in[k], fan_out = N_out[k]:
|
||||
// std::array<float, fan_in*fan_out> weights // row-major
|
||||
// std::array<float, fan_out> biases
|
||||
// std::array<float, fan_out> pre_activation // cached for backprop
|
||||
// std::array<float, fan_out> activation // cached for backprop
|
||||
// std::array<float, fan_in*fan_out> grad_w_accum // for backprop
|
||||
// std::array<float, fan_out> grad_b_accum
|
||||
// * `MLP<NIn, NH1, NH2, NH3, NOut, NMaxExamples, NMaxIterTrain>` — alias
|
||||
// over `MLPCore<FixedStorage<...>>`. All buffers template-sized
|
||||
// std::array, zero heap. This is the firmware model and preserves the
|
||||
// pre-P2 class's exact compile-time surface (`kInput`, `kHidden1..3`,
|
||||
// `kOutput`, `kNumLayers`, `weight_count()` — all constexpr).
|
||||
// * `MLPCore<DynamicStorage>` — runtime-shaped (WASM/native-test/VCV
|
||||
// only; heap at construction time, never per-call). Compile-time
|
||||
// excluded from RP2350 builds.
|
||||
//
|
||||
// Per MLP:
|
||||
// std::array<float, NIn> input_buffer (current set_input values)
|
||||
// std::array<float, NOut> output (post-final-activation; outputs())
|
||||
// std::array<float, NMaxExamples * NIn> dataset_features
|
||||
// std::array<float, NMaxExamples * NOut> dataset_labels
|
||||
// std::size_t dataset_count, dataset_head (FIFO ring buffer)
|
||||
// std::array<float, NMaxIter> loss_history (max iters from train())
|
||||
// Rng rng_
|
||||
// std::array<float, NOut> bp_err_buf, bp_delta_buf (backprop scratch)
|
||||
// BIT-PARITY CONTRACT: for identical shapes and seeds the two storage models
|
||||
// produce bit-identical results — the algorithm code below is shared and
|
||||
// float op order is storage-independent. Enforced by
|
||||
// tests/cpp/test_mlp_storage_parity.cpp.
|
||||
//
|
||||
// FLAT WEIGHT LAYOUT (`get_weights` / `set_weights`)
|
||||
// [layer0_weights ...] [layer1_weights ...] [layer2_weights ...] [layer3_weights ...]
|
||||
// [layer0_biases ...] [layer1_biases ...] [layer2_biases ...] [layer3_biases ...]
|
||||
// Documented in detail near `weight_count()`.
|
||||
//
|
||||
// CONCEPT SATISFACTION
|
||||
// The class satisfies `nisps::MLEngine`:
|
||||
// set_input, process, outputs, add_example, train (no-arg overload
|
||||
// returning float), move_weights(speed, spread), draw_weights(spread),
|
||||
// reset, seed.
|
||||
// Plus diagnostics required by the broader API (see Stream 2 brief in
|
||||
// architecture.md): eval_loss, layer_stats, get/set_weights, weight_count,
|
||||
// infer_batch, loss_history.
|
||||
// The class satisfies `nisps::MLEngine`: set_input, process, outputs,
|
||||
// add_example, train (no-arg overload returning float), move_weights,
|
||||
// draw_weights, reset, seed. Plus diagnostics: eval_loss, layer_stats,
|
||||
// get/set_weights, weight_count, infer_batch, loss_history.
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <array>
|
||||
#include <cmath>
|
||||
#include <cstddef>
|
||||
#include <cstdint>
|
||||
|
|
@ -65,133 +53,36 @@
|
|||
#include "loss.hpp"
|
||||
#include "rl.hpp"
|
||||
#include "stats.hpp"
|
||||
#include "storage.hpp"
|
||||
#include "training.hpp"
|
||||
|
||||
namespace nisps::ml {
|
||||
|
||||
// Layer<FanIn, FanOut, Act>. Stores its weights, biases, and the work
|
||||
// buffers needed for forward + backprop. Header-only, all sizes compile-
|
||||
// time. Each method is small; the compiler will inline through.
|
||||
template <std::size_t FanIn, std::size_t FanOut, Activation Act>
|
||||
struct Layer {
|
||||
static constexpr std::size_t kFanIn = FanIn;
|
||||
static constexpr std::size_t kFanOut = FanOut;
|
||||
static constexpr Activation kAct = Act;
|
||||
// Activation of layer L in the fixed 4-layer topology.
|
||||
template <std::size_t L>
|
||||
inline constexpr Activation kLayerActivation =
|
||||
(L == 3u) ? Activation::Sigmoid : Activation::ReLU;
|
||||
|
||||
std::array<float, FanIn * FanOut> weights{};
|
||||
std::array<float, FanOut> biases{};
|
||||
// Cached during forward(); consumed during backprop().
|
||||
std::array<float, FanOut> pre_activation{};
|
||||
std::array<float, FanOut> activation{};
|
||||
// Gradient accumulators — used per-sample for SGD weight update.
|
||||
std::array<float, FanIn * FanOut> grad_w{};
|
||||
std::array<float, FanOut> grad_b{};
|
||||
|
||||
NISPS_FORCE_INLINE float& w(std::size_t node, std::size_t in) noexcept {
|
||||
return weights[node * FanIn + in];
|
||||
}
|
||||
NISPS_FORCE_INLINE float w(std::size_t node, std::size_t in) const noexcept {
|
||||
return weights[node * FanIn + in];
|
||||
}
|
||||
|
||||
// Forward: compute pre_activation and activation given an input span.
|
||||
NISPS_HOT NISPS_FORCE_INLINE
|
||||
void forward(std::span<const float, FanIn> input) noexcept {
|
||||
for (std::size_t node = 0; node < FanOut; ++node) {
|
||||
const std::size_t row = node * FanIn;
|
||||
float sum = biases[node];
|
||||
for (std::size_t j = 0; j < FanIn; ++j) {
|
||||
sum += weights[row + j] * input[j];
|
||||
}
|
||||
pre_activation[node] = sum;
|
||||
activation[node] = activate<Act>(sum);
|
||||
}
|
||||
}
|
||||
|
||||
// Compute incoming-error vector for the previous layer:
|
||||
// delta_in[j] = sum_node (err_signal[node] * w[node, j])
|
||||
// Where err_signal[node] = upstream_err[node] * d/dpre activation.
|
||||
// Also accumulates per-weight and per-bias gradients (no LR yet).
|
||||
NISPS_HOT NISPS_FORCE_INLINE
|
||||
void backprop_accumulate(std::span<const float, FanIn> input,
|
||||
std::span<const float, FanOut> upstream_err,
|
||||
std::span<float, FanIn> delta_in,
|
||||
float sample_weight) noexcept {
|
||||
for (std::size_t j = 0; j < FanIn; ++j) delta_in[j] = 0.f;
|
||||
|
||||
for (std::size_t node = 0; node < FanOut; ++node) {
|
||||
const float err_signal =
|
||||
upstream_err[node] * activate_deriv_pre<Act>(pre_activation[node]) * sample_weight;
|
||||
const std::size_t row = node * FanIn;
|
||||
for (std::size_t j = 0; j < FanIn; ++j) {
|
||||
grad_w[row + j] += err_signal * input[j];
|
||||
delta_in[j] += err_signal * weights[row + j];
|
||||
}
|
||||
grad_b[node] += err_signal;
|
||||
}
|
||||
}
|
||||
|
||||
// Apply accumulated gradient to weights+biases with clipping. Resets
|
||||
// the accumulators to zero for the next sample/iteration.
|
||||
NISPS_FORCE_INLINE
|
||||
void apply_grad(float lr) noexcept {
|
||||
for (std::size_t i = 0; i < FanIn * FanOut; ++i) {
|
||||
const float g = clip_gradient(grad_w[i]);
|
||||
weights[i] -= lr * g;
|
||||
grad_w[i] = 0.f;
|
||||
}
|
||||
for (std::size_t i = 0; i < FanOut; ++i) {
|
||||
const float g = clip_gradient(grad_b[i]);
|
||||
biases[i] -= lr * g;
|
||||
grad_b[i] = 0.f;
|
||||
}
|
||||
}
|
||||
|
||||
NISPS_FORCE_INLINE
|
||||
void clear_grad() noexcept {
|
||||
for (std::size_t i = 0; i < FanIn * FanOut; ++i) grad_w[i] = 0.f;
|
||||
for (std::size_t i = 0; i < FanOut; ++i) grad_b[i] = 0.f;
|
||||
}
|
||||
};
|
||||
|
||||
// MLP<NIn, NHidden1, NHidden2, NHidden3, NOut, NMaxExamples = 128>
|
||||
//
|
||||
// MaxIterTrain caps the loss-history buffer; if a caller asks for more
|
||||
// iterations they will be honored at runtime, but only the first
|
||||
// kMaxIterTrain are recorded for inspection. 4096 fits the playground's
|
||||
// upper bound and costs 16 KiB.
|
||||
template <std::size_t NIn,
|
||||
std::size_t NHidden1,
|
||||
std::size_t NHidden2,
|
||||
std::size_t NHidden3,
|
||||
std::size_t NOut,
|
||||
std::size_t NMaxExamples = 128u,
|
||||
std::size_t NMaxIterTrain = 4096u>
|
||||
class MLP {
|
||||
template <typename Storage>
|
||||
class MLPCore : public Storage {
|
||||
public:
|
||||
static constexpr std::size_t kInput = NIn;
|
||||
static constexpr std::size_t kHidden1 = NHidden1;
|
||||
static constexpr std::size_t kHidden2 = NHidden2;
|
||||
static constexpr std::size_t kHidden3 = NHidden3;
|
||||
static constexpr std::size_t kOutput = NOut;
|
||||
static constexpr std::size_t kMaxExamples = NMaxExamples;
|
||||
static constexpr std::size_t kMaxIterTrain = NMaxIterTrain;
|
||||
static constexpr std::size_t kNumLayers = 4u;
|
||||
|
||||
using Layer0 = Layer<NIn, NHidden1, Activation::ReLU>;
|
||||
using Layer1 = Layer<NHidden1, NHidden2, Activation::ReLU>;
|
||||
using Layer2 = Layer<NHidden2, NHidden3, Activation::ReLU>;
|
||||
using Layer3 = Layer<NHidden3, NOut, Activation::Sigmoid>;
|
||||
static constexpr std::size_t kNumLayers = kMlpNumLayers;
|
||||
|
||||
// ---------------------------------------------------------------
|
||||
// Lifecycle
|
||||
// Lifecycle. Extra arguments are forwarded to the storage policy —
|
||||
// FixedStorage takes none (`MLP m(seed)`), DynamicStorage takes its
|
||||
// runtime dimensions (`MLPCore<DynamicStorage> m(seed, n_in, hidden,
|
||||
// n_out, ...)`).
|
||||
// ---------------------------------------------------------------
|
||||
explicit MLP(std::uint64_t seed) noexcept : rng_(seed) {
|
||||
template <typename... StorageArgs>
|
||||
explicit MLPCore(std::uint64_t seed, StorageArgs&&... storage_args) noexcept
|
||||
: Storage(static_cast<StorageArgs&&>(storage_args)...), rng_(seed) {
|
||||
// Default-init weights with spread=1 (Xavier-like). The IML
|
||||
// interface caller is expected to draw_weights() with the
|
||||
// playground spread before the first inference; this default
|
||||
// simply gives us a non-degenerate starting state for tests
|
||||
// that skip an explicit draw.
|
||||
// interface caller is expected to draw_weights() with its own
|
||||
// spread before the first inference; this default simply gives a
|
||||
// non-degenerate starting state for tests that skip an explicit
|
||||
// draw.
|
||||
if (!storage_ok_()) return;
|
||||
draw_weights(1.f);
|
||||
clear_dataset_();
|
||||
loss_history_count_ = 0u;
|
||||
|
|
@ -201,48 +92,54 @@ class MLP {
|
|||
// Inference API (concept: set_input / process / outputs)
|
||||
// ---------------------------------------------------------------
|
||||
NISPS_FORCE_INLINE void set_input(std::size_t i, float v) noexcept {
|
||||
if (i < NIn) input_[i] = v;
|
||||
if (!storage_ok_()) return;
|
||||
if (i < this->n_in()) this->input_buf()[i] = v;
|
||||
}
|
||||
|
||||
NISPS_HOT void process() noexcept {
|
||||
forward_(std::span<const float, NIn>(input_));
|
||||
if (!storage_ok_()) return;
|
||||
forward_(this->input_buf());
|
||||
// Mirror final activation into the output buffer so callers can
|
||||
// read a stable span.
|
||||
const auto& a = layer3_.activation;
|
||||
for (std::size_t i = 0; i < NOut; ++i) output_[i] = a[i];
|
||||
const auto a = this->template act_l<3u>();
|
||||
auto out = this->output_buf();
|
||||
const std::size_t n_out = this->n_out();
|
||||
for (std::size_t i = 0; i < n_out; ++i) out[i] = a[i];
|
||||
}
|
||||
|
||||
NISPS_FORCE_INLINE std::span<const float> outputs() const noexcept {
|
||||
return std::span<const float>(output_.data(), NOut);
|
||||
return this->output_buf();
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------
|
||||
// Dataset / Training (concept: add_example / train)
|
||||
// ---------------------------------------------------------------
|
||||
// FIFO ring buffer; oldest example evicted when full. No allocation.
|
||||
// We don't track logical insertion order during training because SGD
|
||||
// doesn't care — the iteration order over the buffer is arbitrary.
|
||||
void add_example(std::span<const float> features,
|
||||
std::span<const float> labels) noexcept {
|
||||
if (features.size() < NIn || labels.size() < NOut) return;
|
||||
if (!storage_ok_()) return;
|
||||
const std::size_t n_in = this->n_in();
|
||||
const std::size_t n_out = this->n_out();
|
||||
if (features.size() < n_in || labels.size() < n_out) return;
|
||||
|
||||
std::size_t slot;
|
||||
if (dataset_count_ < NMaxExamples) {
|
||||
if (dataset_count_ < this->max_examples()) {
|
||||
slot = dataset_count_++;
|
||||
} else {
|
||||
// Buffer full: overwrite the slot pointed at by head_ (oldest)
|
||||
// and advance head_ to the next-oldest.
|
||||
slot = dataset_head_;
|
||||
dataset_head_ = (dataset_head_ + 1u) % NMaxExamples;
|
||||
dataset_head_ = (dataset_head_ + 1u) % this->max_examples();
|
||||
}
|
||||
const std::size_t f_off = slot * NIn;
|
||||
const std::size_t l_off = slot * NOut;
|
||||
for (std::size_t i = 0; i < NIn; ++i) ds_features_[f_off + i] = features[i];
|
||||
for (std::size_t i = 0; i < NOut; ++i) ds_labels_ [l_off + i] = labels[i];
|
||||
auto dsf = this->ds_features();
|
||||
auto dsl = this->ds_labels();
|
||||
const std::size_t f_off = slot * n_in;
|
||||
const std::size_t l_off = slot * n_out;
|
||||
for (std::size_t i = 0; i < n_in; ++i) dsf[f_off + i] = features[i];
|
||||
for (std::size_t i = 0; i < n_out; ++i) dsl[l_off + i] = labels[i];
|
||||
}
|
||||
|
||||
// Concept-required no-arg overload. Default learning rate matches the
|
||||
// playground's "sane RL training" knob; max_iter and min_err follow.
|
||||
// Concept-required no-arg overload.
|
||||
float train() noexcept {
|
||||
return train(1.f, 1000u, 0.001f, std::span<const float>{});
|
||||
}
|
||||
|
|
@ -251,18 +148,21 @@ class MLP {
|
|||
// current example count and sum to 1.0 (caller's responsibility — we
|
||||
// do NOT renormalize).
|
||||
//
|
||||
// Returns final epoch loss. Records per-iteration loss in
|
||||
// `loss_history_` (bounded by kMaxIterTrain).
|
||||
// Returns final epoch loss. Records per-iteration loss in the loss
|
||||
// history (bounded by max_iter_train()).
|
||||
float train(float lr,
|
||||
std::size_t max_iter,
|
||||
float min_err,
|
||||
std::span<const float> sample_weights = {}) noexcept {
|
||||
loss_history_count_ = 0u;
|
||||
if (!storage_ok_()) return 0.f;
|
||||
if (dataset_count_ == 0u) return 0.f;
|
||||
|
||||
const bool weighted = !sample_weights.empty();
|
||||
const bool weighted = !sample_weights.empty();
|
||||
const float uniform_w = 1.f / static_cast<float>(dataset_count_);
|
||||
|
||||
auto loss_hist = this->loss_hist_buf();
|
||||
|
||||
float epoch_loss = 0.f;
|
||||
for (std::size_t iter = 0; iter < max_iter; ++iter) {
|
||||
epoch_loss = 0.f;
|
||||
|
|
@ -274,31 +174,34 @@ class MLP {
|
|||
const float w = weighted ? sample_weights[s] : uniform_w;
|
||||
|
||||
// Forward pass on sample s.
|
||||
std::span<const float, NIn> x = sample_features_(s);
|
||||
std::span<const float> x = sample_features_(s);
|
||||
forward_(x);
|
||||
|
||||
// Per-sample loss (NOT scaled by 1/N — the meml-ues fix).
|
||||
std::array<float, NOut> deriv{};
|
||||
// The eval scratch of the final layer doubles as the loss-
|
||||
// derivative buffer (mse_per_sample fully overwrites it;
|
||||
// eval_loss never runs concurrently).
|
||||
auto deriv = this->template eval_act_l<3u>();
|
||||
const float sample_loss = mse_per_sample(
|
||||
sample_labels_(s),
|
||||
std::span<const float>(layer3_.activation.data(), NOut),
|
||||
std::span<float>(deriv.data(), NOut));
|
||||
std::span<const float>(this->template act_l<3u>()),
|
||||
deriv);
|
||||
|
||||
// Aggregate weighted loss.
|
||||
epoch_loss += w * sample_loss;
|
||||
|
||||
// Backprop with the same w as the gradient scaler.
|
||||
backprop_(x, std::span<const float, NOut>(deriv), w);
|
||||
backprop_(x, deriv, w);
|
||||
|
||||
// Apply gradient (per-sample, SGD).
|
||||
layer3_.apply_grad(lr);
|
||||
layer2_.apply_grad(lr);
|
||||
layer1_.apply_grad(lr);
|
||||
layer0_.apply_grad(lr);
|
||||
apply_grad_<3u>(lr);
|
||||
apply_grad_<2u>(lr);
|
||||
apply_grad_<1u>(lr);
|
||||
apply_grad_<0u>(lr);
|
||||
}
|
||||
|
||||
if (loss_history_count_ < NMaxIterTrain) {
|
||||
loss_history_[loss_history_count_++] = epoch_loss;
|
||||
if (loss_history_count_ < this->max_iter_train()) {
|
||||
loss_hist[loss_history_count_++] = epoch_loss;
|
||||
}
|
||||
|
||||
if (epoch_loss < min_err) break;
|
||||
|
|
@ -311,40 +214,46 @@ class MLP {
|
|||
// ---------------------------------------------------------------
|
||||
void move_weights(float speed, float spread,
|
||||
std::span<const std::uint8_t> output_pin_mask = {}) noexcept {
|
||||
move_weights_layer(std::span<float>(layer0_.weights), std::span<float>(layer0_.biases),
|
||||
Layer0::kFanIn, speed, spread, /*final=*/false, {}, rng_);
|
||||
move_weights_layer(std::span<float>(layer1_.weights), std::span<float>(layer1_.biases),
|
||||
Layer1::kFanIn, speed, spread, /*final=*/false, {}, rng_);
|
||||
move_weights_layer(std::span<float>(layer2_.weights), std::span<float>(layer2_.biases),
|
||||
Layer2::kFanIn, speed, spread, /*final=*/false, {}, rng_);
|
||||
move_weights_layer(std::span<float>(layer3_.weights), std::span<float>(layer3_.biases),
|
||||
Layer3::kFanIn, speed, spread, /*final=*/true, output_pin_mask, rng_);
|
||||
if (!storage_ok_()) return;
|
||||
move_weights_layer(this->template weights_l<0u>(), this->template biases_l<0u>(),
|
||||
this->template fan_in_l<0u>(), speed, spread, /*final=*/false, {}, rng_);
|
||||
move_weights_layer(this->template weights_l<1u>(), this->template biases_l<1u>(),
|
||||
this->template fan_in_l<1u>(), speed, spread, /*final=*/false, {}, rng_);
|
||||
move_weights_layer(this->template weights_l<2u>(), this->template biases_l<2u>(),
|
||||
this->template fan_in_l<2u>(), speed, spread, /*final=*/false, {}, rng_);
|
||||
move_weights_layer(this->template weights_l<3u>(), this->template biases_l<3u>(),
|
||||
this->template fan_in_l<3u>(), speed, spread, /*final=*/true,
|
||||
output_pin_mask, rng_);
|
||||
}
|
||||
|
||||
void draw_weights(float spread) noexcept {
|
||||
draw_weights_layer(std::span<float>(layer0_.weights), std::span<float>(layer0_.biases),
|
||||
Layer0::kFanIn, spread, rng_);
|
||||
draw_weights_layer(std::span<float>(layer1_.weights), std::span<float>(layer1_.biases),
|
||||
Layer1::kFanIn, spread, rng_);
|
||||
draw_weights_layer(std::span<float>(layer2_.weights), std::span<float>(layer2_.biases),
|
||||
Layer2::kFanIn, spread, rng_);
|
||||
draw_weights_layer(std::span<float>(layer3_.weights), std::span<float>(layer3_.biases),
|
||||
Layer3::kFanIn, spread, rng_);
|
||||
layer0_.clear_grad();
|
||||
layer1_.clear_grad();
|
||||
layer2_.clear_grad();
|
||||
layer3_.clear_grad();
|
||||
if (!storage_ok_()) return;
|
||||
draw_weights_layer(this->template weights_l<0u>(), this->template biases_l<0u>(),
|
||||
this->template fan_in_l<0u>(), spread, rng_);
|
||||
draw_weights_layer(this->template weights_l<1u>(), this->template biases_l<1u>(),
|
||||
this->template fan_in_l<1u>(), spread, rng_);
|
||||
draw_weights_layer(this->template weights_l<2u>(), this->template biases_l<2u>(),
|
||||
this->template fan_in_l<2u>(), spread, rng_);
|
||||
draw_weights_layer(this->template weights_l<3u>(), this->template biases_l<3u>(),
|
||||
this->template fan_in_l<3u>(), spread, rng_);
|
||||
clear_grad_<0u>();
|
||||
clear_grad_<1u>();
|
||||
clear_grad_<2u>();
|
||||
clear_grad_<3u>();
|
||||
}
|
||||
|
||||
// Concept reset: clear weights, dataset, and loss history. Seed is
|
||||
// intentionally NOT reset (use `seed()` for that).
|
||||
void reset() noexcept {
|
||||
if (!storage_ok_()) return;
|
||||
clear_dataset_();
|
||||
loss_history_count_ = 0u;
|
||||
// Re-init weights from current rng state with default spread.
|
||||
draw_weights(1.f);
|
||||
for (std::size_t i = 0; i < NIn; ++i) input_[i] = 0.f;
|
||||
for (std::size_t i = 0; i < NOut; ++i) output_[i] = 0.f;
|
||||
auto in = this->input_buf();
|
||||
auto out = this->output_buf();
|
||||
for (std::size_t i = 0; i < in.size(); ++i) in[i] = 0.f;
|
||||
for (std::size_t i = 0; i < out.size(); ++i) out[i] = 0.f;
|
||||
}
|
||||
|
||||
void seed(std::uint64_t s) noexcept { rng_.seed(s); }
|
||||
|
|
@ -352,41 +261,28 @@ class MLP {
|
|||
// ---------------------------------------------------------------
|
||||
// Diagnostics
|
||||
// ---------------------------------------------------------------
|
||||
// Forward pass + MSE on a single (input, label) implied by current
|
||||
// input_ and the most recent training labels — i.e. "what would the
|
||||
// loss be on the current input if the label were the current output?"
|
||||
// For now we report the average loss across the training set without
|
||||
// updating weights. Useful for non-destructive evaluation.
|
||||
// Average MSE across the training set without updating weights or the
|
||||
// cached activations (runs through the mutable eval scratch).
|
||||
float eval_loss() const noexcept {
|
||||
if (!storage_ok_()) return 0.f;
|
||||
if (dataset_count_ == 0u) return 0.f;
|
||||
const float inv_n = 1.f / static_cast<float>(dataset_count_);
|
||||
// We need a non-const forward pass to use the cached buffers; since
|
||||
// eval_loss is logically const, fork a local computation that
|
||||
// doesn't touch member buffers. That means recomputing through the
|
||||
// layer weights against scratch arrays — no allocation, just stack.
|
||||
const std::size_t n_out = this->n_out();
|
||||
auto dsl = this->ds_labels();
|
||||
|
||||
float total = 0.f;
|
||||
for (std::size_t s = 0; s < dataset_count_; ++s) {
|
||||
std::array<float, NHidden1> a1{};
|
||||
std::array<float, NHidden2> a2{};
|
||||
std::array<float, NHidden3> a3{};
|
||||
std::array<float, NOut> ao{};
|
||||
|
||||
const std::size_t f_off = s * NIn;
|
||||
forward_const_layer<NIn, NHidden1, Activation::ReLU>(
|
||||
std::span<const float, NIn>(ds_features_.data() + f_off, NIn),
|
||||
layer0_.weights, layer0_.biases, a1);
|
||||
forward_const_layer<NHidden1, NHidden2, Activation::ReLU>(
|
||||
std::span<const float, NHidden1>(a1), layer1_.weights, layer1_.biases, a2);
|
||||
forward_const_layer<NHidden2, NHidden3, Activation::ReLU>(
|
||||
std::span<const float, NHidden2>(a2), layer2_.weights, layer2_.biases, a3);
|
||||
forward_const_layer<NHidden3, NOut, Activation::Sigmoid>(
|
||||
std::span<const float, NHidden3>(a3), layer3_.weights, layer3_.biases, ao);
|
||||
forward_eval_layer_<0u>(sample_features_(s));
|
||||
forward_eval_layer_<1u>(this->template eval_act_l<0u>());
|
||||
forward_eval_layer_<2u>(this->template eval_act_l<1u>());
|
||||
forward_eval_layer_<3u>(this->template eval_act_l<2u>());
|
||||
|
||||
const auto ao = this->template eval_act_l<3u>();
|
||||
float sse = 0.f;
|
||||
const float inv_o = 1.f / static_cast<float>(NOut);
|
||||
const std::size_t l_off = s * NOut;
|
||||
for (std::size_t j = 0; j < NOut; ++j) {
|
||||
const float d = ds_labels_[l_off + j] - ao[j];
|
||||
const float inv_o = 1.f / static_cast<float>(n_out);
|
||||
const std::size_t l_off = s * n_out;
|
||||
for (std::size_t j = 0; j < n_out; ++j) {
|
||||
const float d = dsl[l_off + j] - ao[j];
|
||||
sse += d * d * inv_o;
|
||||
}
|
||||
total += sse * inv_n;
|
||||
|
|
@ -395,73 +291,77 @@ class MLP {
|
|||
}
|
||||
|
||||
LayerStats layer_stats(std::size_t layer_idx) const noexcept {
|
||||
if (!storage_ok_()) return {};
|
||||
switch (layer_idx) {
|
||||
case 0: return compute_layer_stats(layer0_.weights, layer0_.biases);
|
||||
case 1: return compute_layer_stats(layer1_.weights, layer1_.biases);
|
||||
case 2: return compute_layer_stats(layer2_.weights, layer2_.biases);
|
||||
case 3: return compute_layer_stats(layer3_.weights, layer3_.biases);
|
||||
case 0: return compute_layer_stats(this->template weights_l<0u>(),
|
||||
this->template biases_l<0u>());
|
||||
case 1: return compute_layer_stats(this->template weights_l<1u>(),
|
||||
this->template biases_l<1u>());
|
||||
case 2: return compute_layer_stats(this->template weights_l<2u>(),
|
||||
this->template biases_l<2u>());
|
||||
case 3: return compute_layer_stats(this->template weights_l<3u>(),
|
||||
this->template biases_l<3u>());
|
||||
default: return {};
|
||||
}
|
||||
}
|
||||
|
||||
// Flat layout: layer0 weights, layer1 weights, layer2 weights, layer3
|
||||
// weights, then layer0..3 biases. The split is `weights first all
|
||||
// layers, then biases all layers` so callers serializing weights can
|
||||
// pre-compute offsets without consulting layer-specific tables.
|
||||
static constexpr std::size_t weight_count() noexcept {
|
||||
return NIn * NHidden1 + NHidden1 * NHidden2 + NHidden2 * NHidden3 + NHidden3 * NOut
|
||||
+ NHidden1 + NHidden2 + NHidden3 + NOut;
|
||||
}
|
||||
|
||||
// Returns a span into a member-owned scratch buffer that holds a copy
|
||||
// Returns a span into a storage-owned scratch buffer that holds a copy
|
||||
// of the flat weights+biases. The buffer is regenerated on each call,
|
||||
// so don't hold onto the span across mutations.
|
||||
std::span<const float> get_weights() noexcept {
|
||||
if (!storage_ok_()) return {};
|
||||
auto flat = this->flat_buf();
|
||||
std::size_t k = 0u;
|
||||
// Weights, layer-major.
|
||||
for (float v : layer0_.weights) flat_weight_buf_[k++] = v;
|
||||
for (float v : layer1_.weights) flat_weight_buf_[k++] = v;
|
||||
for (float v : layer2_.weights) flat_weight_buf_[k++] = v;
|
||||
for (float v : layer3_.weights) flat_weight_buf_[k++] = v;
|
||||
for (float v : this->template weights_l<0u>()) flat[k++] = v;
|
||||
for (float v : this->template weights_l<1u>()) flat[k++] = v;
|
||||
for (float v : this->template weights_l<2u>()) flat[k++] = v;
|
||||
for (float v : this->template weights_l<3u>()) flat[k++] = v;
|
||||
// Biases.
|
||||
for (float v : layer0_.biases) flat_weight_buf_[k++] = v;
|
||||
for (float v : layer1_.biases) flat_weight_buf_[k++] = v;
|
||||
for (float v : layer2_.biases) flat_weight_buf_[k++] = v;
|
||||
for (float v : layer3_.biases) flat_weight_buf_[k++] = v;
|
||||
return std::span<const float>(flat_weight_buf_.data(), weight_count());
|
||||
for (float v : this->template biases_l<0u>()) flat[k++] = v;
|
||||
for (float v : this->template biases_l<1u>()) flat[k++] = v;
|
||||
for (float v : this->template biases_l<2u>()) flat[k++] = v;
|
||||
for (float v : this->template biases_l<3u>()) flat[k++] = v;
|
||||
return std::span<const float>(flat.data(), k);
|
||||
}
|
||||
|
||||
void set_weights(std::span<const float> w) noexcept {
|
||||
if (w.size() < weight_count()) return;
|
||||
if (!storage_ok_()) return;
|
||||
if (w.size() < this->weight_count()) return;
|
||||
std::size_t k = 0u;
|
||||
for (float& v : layer0_.weights) v = w[k++];
|
||||
for (float& v : layer1_.weights) v = w[k++];
|
||||
for (float& v : layer2_.weights) v = w[k++];
|
||||
for (float& v : layer3_.weights) v = w[k++];
|
||||
for (float& v : layer0_.biases) v = w[k++];
|
||||
for (float& v : layer1_.biases) v = w[k++];
|
||||
for (float& v : layer2_.biases) v = w[k++];
|
||||
for (float& v : layer3_.biases) v = w[k++];
|
||||
for (float& v : this->template weights_l<0u>()) v = w[k++];
|
||||
for (float& v : this->template weights_l<1u>()) v = w[k++];
|
||||
for (float& v : this->template weights_l<2u>()) v = w[k++];
|
||||
for (float& v : this->template weights_l<3u>()) v = w[k++];
|
||||
for (float& v : this->template biases_l<0u>()) v = w[k++];
|
||||
for (float& v : this->template biases_l<1u>()) v = w[k++];
|
||||
for (float& v : this->template biases_l<2u>()) v = w[k++];
|
||||
for (float& v : this->template biases_l<3u>()) v = w[k++];
|
||||
}
|
||||
|
||||
// Run inference on N points (each NIn-sized) and write N output vectors
|
||||
// (each NOut-sized) into `outs`. NO heap. Modifies the internal cached
|
||||
// Run inference on N points (each n_in-sized) and write N output vectors
|
||||
// (each n_out-sized) into `outs`. NO heap. Modifies the internal cached
|
||||
// activations as a side effect.
|
||||
void infer_batch(std::span<const float> points,
|
||||
std::span<float> outs) noexcept {
|
||||
const std::size_t n = points.size() / NIn;
|
||||
if (outs.size() < n * NOut) return;
|
||||
if (!storage_ok_()) return;
|
||||
const std::size_t n_in = this->n_in();
|
||||
const std::size_t n_out = this->n_out();
|
||||
const std::size_t n = points.size() / n_in;
|
||||
if (outs.size() < n * n_out) return;
|
||||
auto in = this->input_buf();
|
||||
auto out = this->output_buf();
|
||||
for (std::size_t i = 0; i < n; ++i) {
|
||||
const std::size_t in_off = i * NIn;
|
||||
for (std::size_t j = 0; j < NIn; ++j) input_[j] = points[in_off + j];
|
||||
const std::size_t in_off = i * n_in;
|
||||
for (std::size_t j = 0; j < n_in; ++j) in[j] = points[in_off + j];
|
||||
process();
|
||||
const std::size_t out_off = i * NOut;
|
||||
for (std::size_t j = 0; j < NOut; ++j) outs[out_off + j] = output_[j];
|
||||
const std::size_t out_off = i * n_out;
|
||||
for (std::size_t j = 0; j < n_out; ++j) outs[out_off + j] = out[j];
|
||||
}
|
||||
}
|
||||
|
||||
std::span<const float> loss_history() const noexcept {
|
||||
return std::span<const float>(loss_history_.data(), loss_history_count_);
|
||||
return std::span<const float>(this->loss_hist_buf().data(), loss_history_count_);
|
||||
}
|
||||
|
||||
std::size_t example_count() const noexcept { return dataset_count_; }
|
||||
|
|
@ -472,68 +372,141 @@ class MLP {
|
|||
// ---------------------------------------------------------------
|
||||
// Internal helpers
|
||||
// ---------------------------------------------------------------
|
||||
// DynamicStorage construction can fail (arena allocation); FixedStorage
|
||||
// cannot. The check is compile-time `true` for storages without a
|
||||
// `valid()` member, so the fixed/firmware path carries no branch.
|
||||
NISPS_FORCE_INLINE bool storage_ok_() const noexcept {
|
||||
if constexpr (requires(const Storage& s) { { s.valid() } -> std::convertible_to<bool>; }) {
|
||||
return this->valid();
|
||||
} else {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
|
||||
template <std::size_t L>
|
||||
NISPS_HOT NISPS_FORCE_INLINE void forward_layer_(std::span<const float> in) noexcept {
|
||||
const std::size_t fan_in = this->template fan_in_l<L>();
|
||||
const std::size_t fan_out = this->template fan_out_l<L>();
|
||||
auto w = this->template weights_l<L>();
|
||||
auto b = this->template biases_l<L>();
|
||||
auto pa = this->template pre_act_l<L>();
|
||||
auto a = this->template act_l<L>();
|
||||
for (std::size_t node = 0; node < fan_out; ++node) {
|
||||
const std::size_t row = node * fan_in;
|
||||
float sum = b[node];
|
||||
for (std::size_t j = 0; j < fan_in; ++j) {
|
||||
sum += w[row + j] * in[j];
|
||||
}
|
||||
pa[node] = sum;
|
||||
a[node] = activate<kLayerActivation<L>>(sum);
|
||||
}
|
||||
}
|
||||
|
||||
NISPS_HOT NISPS_FORCE_INLINE
|
||||
void forward_(std::span<const float, NIn> in) noexcept {
|
||||
layer0_.forward(in);
|
||||
layer1_.forward(std::span<const float, NHidden1>(layer0_.activation));
|
||||
layer2_.forward(std::span<const float, NHidden2>(layer1_.activation));
|
||||
layer3_.forward(std::span<const float, NHidden3>(layer2_.activation));
|
||||
void forward_(std::span<const float> in) noexcept {
|
||||
forward_layer_<0u>(in);
|
||||
forward_layer_<1u>(this->template act_l<0u>());
|
||||
forward_layer_<2u>(this->template act_l<1u>());
|
||||
forward_layer_<3u>(this->template act_l<2u>());
|
||||
}
|
||||
|
||||
// Backprop one layer: compute the incoming-error vector for the previous
|
||||
// layer into delta_l<L>() and accumulate per-weight/per-bias gradients.
|
||||
template <std::size_t L>
|
||||
NISPS_HOT NISPS_FORCE_INLINE
|
||||
void backprop_layer_(std::span<const float> input,
|
||||
std::span<const float> upstream_err,
|
||||
float sample_weight) noexcept {
|
||||
const std::size_t fan_in = this->template fan_in_l<L>();
|
||||
const std::size_t fan_out = this->template fan_out_l<L>();
|
||||
auto w = this->template weights_l<L>();
|
||||
auto pa = this->template pre_act_l<L>();
|
||||
auto gw = this->template grad_w_l<L>();
|
||||
auto gb = this->template grad_b_l<L>();
|
||||
auto delta_in = this->template delta_l<L>();
|
||||
|
||||
for (std::size_t j = 0; j < fan_in; ++j) delta_in[j] = 0.f;
|
||||
|
||||
for (std::size_t node = 0; node < fan_out; ++node) {
|
||||
const float err_signal =
|
||||
upstream_err[node] *
|
||||
activate_deriv_pre<kLayerActivation<L>>(pa[node]) * sample_weight;
|
||||
const std::size_t row = node * fan_in;
|
||||
for (std::size_t j = 0; j < fan_in; ++j) {
|
||||
gw[row + j] += err_signal * input[j];
|
||||
delta_in[j] += err_signal * w[row + j];
|
||||
}
|
||||
gb[node] += err_signal;
|
||||
}
|
||||
}
|
||||
|
||||
// Backprop with sample_weight applied to every error signal (so the
|
||||
// accumulated gradient is already weighted). No weight update happens
|
||||
// here — caller does it after each sample.
|
||||
NISPS_HOT NISPS_FORCE_INLINE
|
||||
void backprop_(std::span<const float, NIn> input,
|
||||
std::span<const float, NOut> output_deriv,
|
||||
void backprop_(std::span<const float> input,
|
||||
std::span<const float> output_deriv,
|
||||
float sample_weight) noexcept {
|
||||
std::array<float, NHidden3> d3{};
|
||||
std::array<float, NHidden2> d2{};
|
||||
std::array<float, NHidden1> d1{};
|
||||
std::array<float, NIn> d0{};
|
||||
|
||||
layer3_.backprop_accumulate(
|
||||
std::span<const float, NHidden3>(layer2_.activation),
|
||||
output_deriv,
|
||||
std::span<float, NHidden3>(d3),
|
||||
sample_weight);
|
||||
layer2_.backprop_accumulate(
|
||||
std::span<const float, NHidden2>(layer1_.activation),
|
||||
std::span<const float, NHidden3>(d3),
|
||||
std::span<float, NHidden2>(d2),
|
||||
1.f); // weight already in d3
|
||||
layer1_.backprop_accumulate(
|
||||
std::span<const float, NHidden1>(layer0_.activation),
|
||||
std::span<const float, NHidden2>(d2),
|
||||
std::span<float, NHidden1>(d1),
|
||||
1.f);
|
||||
layer0_.backprop_accumulate(
|
||||
input,
|
||||
std::span<const float, NHidden1>(d1),
|
||||
std::span<float, NIn>(d0),
|
||||
1.f);
|
||||
backprop_layer_<3u>(this->template act_l<2u>(), output_deriv, sample_weight);
|
||||
backprop_layer_<2u>(this->template act_l<1u>(), this->template delta_l<3u>(), 1.f);
|
||||
backprop_layer_<1u>(this->template act_l<0u>(), this->template delta_l<2u>(), 1.f);
|
||||
backprop_layer_<0u>(input, this->template delta_l<1u>(), 1.f);
|
||||
}
|
||||
|
||||
// const forward pass for diagnostics. Doesn't touch layer caches.
|
||||
template <std::size_t Fi, std::size_t Fo, Activation A>
|
||||
static NISPS_FORCE_INLINE void forward_const_layer(
|
||||
std::span<const float, Fi> in,
|
||||
const std::array<float, Fi*Fo>& w,
|
||||
const std::array<float, Fo>& b,
|
||||
std::array<float, Fo>& out) noexcept {
|
||||
for (std::size_t node = 0; node < Fo; ++node) {
|
||||
const std::size_t row = node * Fi;
|
||||
float sum = b[node];
|
||||
for (std::size_t j = 0; j < Fi; ++j) sum += w[row + j] * in[j];
|
||||
out[node] = activate<A>(sum);
|
||||
// Apply accumulated gradient to weights+biases with clipping. Resets
|
||||
// the accumulators to zero for the next sample/iteration.
|
||||
template <std::size_t L>
|
||||
NISPS_FORCE_INLINE void apply_grad_(float lr) noexcept {
|
||||
auto w = this->template weights_l<L>();
|
||||
auto b = this->template biases_l<L>();
|
||||
auto gw = this->template grad_w_l<L>();
|
||||
auto gb = this->template grad_b_l<L>();
|
||||
const std::size_t nw = gw.size();
|
||||
const std::size_t nb = gb.size();
|
||||
for (std::size_t i = 0; i < nw; ++i) {
|
||||
const float g = clip_gradient(gw[i]);
|
||||
w[i] -= lr * g;
|
||||
gw[i] = 0.f;
|
||||
}
|
||||
for (std::size_t i = 0; i < nb; ++i) {
|
||||
const float g = clip_gradient(gb[i]);
|
||||
b[i] -= lr * g;
|
||||
gb[i] = 0.f;
|
||||
}
|
||||
}
|
||||
|
||||
NISPS_FORCE_INLINE std::span<const float, NIn> sample_features_(std::size_t s) const noexcept {
|
||||
return std::span<const float, NIn>(ds_features_.data() + s * NIn, NIn);
|
||||
template <std::size_t L>
|
||||
NISPS_FORCE_INLINE void clear_grad_() noexcept {
|
||||
auto gw = this->template grad_w_l<L>();
|
||||
auto gb = this->template grad_b_l<L>();
|
||||
for (std::size_t i = 0; i < gw.size(); ++i) gw[i] = 0.f;
|
||||
for (std::size_t i = 0; i < gb.size(); ++i) gb[i] = 0.f;
|
||||
}
|
||||
|
||||
// const forward pass for diagnostics — writes into the mutable eval
|
||||
// scratch, never the real caches.
|
||||
template <std::size_t L>
|
||||
NISPS_FORCE_INLINE void forward_eval_layer_(std::span<const float> in) const noexcept {
|
||||
const std::size_t fan_in = this->template fan_in_l<L>();
|
||||
const std::size_t fan_out = this->template fan_out_l<L>();
|
||||
auto w = this->template weights_l<L>();
|
||||
auto b = this->template biases_l<L>();
|
||||
auto out = this->template eval_act_l<L>();
|
||||
for (std::size_t node = 0; node < fan_out; ++node) {
|
||||
const std::size_t row = node * fan_in;
|
||||
float sum = b[node];
|
||||
for (std::size_t j = 0; j < fan_in; ++j) sum += w[row + j] * in[j];
|
||||
out[node] = activate<kLayerActivation<L>>(sum);
|
||||
}
|
||||
}
|
||||
|
||||
NISPS_FORCE_INLINE std::span<const float> sample_features_(std::size_t s) const noexcept {
|
||||
const std::size_t n_in = this->n_in();
|
||||
return this->ds_features().subspan(s * n_in, n_in);
|
||||
}
|
||||
NISPS_FORCE_INLINE std::span<const float> sample_labels_(std::size_t s) const noexcept {
|
||||
return std::span<const float>(ds_labels_.data() + s * NOut, NOut);
|
||||
const std::size_t n_out = this->n_out();
|
||||
return this->ds_labels().subspan(s * n_out, n_out);
|
||||
}
|
||||
|
||||
void clear_dataset_() noexcept {
|
||||
|
|
@ -542,28 +515,28 @@ class MLP {
|
|||
}
|
||||
|
||||
// ---------------------------------------------------------------
|
||||
// Members
|
||||
// Members (shape-independent; everything sized lives in Storage)
|
||||
// ---------------------------------------------------------------
|
||||
Layer0 layer0_{};
|
||||
Layer1 layer1_{};
|
||||
Layer2 layer2_{};
|
||||
Layer3 layer3_{};
|
||||
|
||||
std::array<float, NIn> input_{};
|
||||
std::array<float, NOut> output_{};
|
||||
|
||||
std::array<float, NMaxExamples * NIn> ds_features_{};
|
||||
std::array<float, NMaxExamples * NOut> ds_labels_{};
|
||||
std::size_t dataset_count_ = 0u;
|
||||
std::size_t dataset_head_ = 0u;
|
||||
|
||||
std::array<float, weight_count()> flat_weight_buf_{};
|
||||
std::array<float, NMaxIterTrain> loss_history_{};
|
||||
std::size_t dataset_count_ = 0u;
|
||||
std::size_t dataset_head_ = 0u;
|
||||
std::size_t loss_history_count_ = 0u;
|
||||
|
||||
Rng rng_;
|
||||
};
|
||||
|
||||
// The classic fixed-architecture MLP — the firmware model and the default
|
||||
// everywhere a compile-time shape is known. `MLP<NIn, 10, 10, 14, NOut>`
|
||||
// matches the legacy firmware default [10, 10, 14].
|
||||
template <std::size_t NIn,
|
||||
std::size_t NHidden1,
|
||||
std::size_t NHidden2,
|
||||
std::size_t NHidden3,
|
||||
std::size_t NOut,
|
||||
std::size_t NMaxExamples = 128u,
|
||||
std::size_t NMaxIterTrain = 4096u>
|
||||
using MLP = MLPCore<
|
||||
FixedStorage<NIn, NHidden1, NHidden2, NHidden3, NOut, NMaxExamples, NMaxIterTrain>>;
|
||||
|
||||
// MLP satisfies the MLEngine concept. We keep a static_assert in the test
|
||||
// suite (test_mlp_concept_satisfied) — see test_mlp_init.cpp.
|
||||
|
||||
|
|
|
|||
203
nisps/ml/storage.hpp
Normal file
203
nisps/ml/storage.hpp
Normal file
|
|
@ -0,0 +1,203 @@
|
|||
// nisps/ml/storage.hpp — storage policies for the MLP core (fixed flavour).
|
||||
//
|
||||
// The MLP algorithms (nisps/ml/mlp.hpp `MLPCore<Storage>`) are written ONCE
|
||||
// against a storage concept; the storage supplies every dimension and every
|
||||
// buffer. Two models exist:
|
||||
//
|
||||
// * `FixedStorage<NIn, NH1, NH2, NH3, NOut, NMaxExamples, NMaxIterTrain>`
|
||||
// (this file) — all buffers are template-sized `std::array`, zero heap,
|
||||
// `NISPS_AUDIO_MEM`-able. This is the firmware model; the classic
|
||||
// `MLP<...>` template is an alias over it and its compile-time constants
|
||||
// (`kInput`, `kHidden1..3`, `kOutput`, `weight_count()`) are preserved.
|
||||
//
|
||||
// * `DynamicStorage` (nisps/ml/dynamic_storage.hpp) — dimensions chosen at
|
||||
// construction, one arena allocation, no allocation after construction.
|
||||
// Compile-time EXCLUDED from embedded builds (see NISPS_TARGET_EMBEDDED
|
||||
// in nisps/core/perf.hpp).
|
||||
//
|
||||
// STORAGE SURFACE (both models; L is the layer index 0..3)
|
||||
// dims: n_in(), n_out(), fan_in_l<L>(), fan_out_l<L>(),
|
||||
// max_examples(), max_iter_train(), weight_count()
|
||||
// layers: weights_l<L>(), biases_l<L>(), pre_act_l<L>(), act_l<L>(),
|
||||
// grad_w_l<L>(), grad_b_l<L>(), delta_l<L>() [backprop scratch,
|
||||
// sized fan_in(L)], eval_act_l<L>() [const-eval scratch,
|
||||
// sized fan_out(L), mutable]
|
||||
// global: input_buf(), output_buf(), ds_features(), ds_labels(),
|
||||
// flat_buf(), loss_hist_buf()
|
||||
//
|
||||
// For `FixedStorage` every dim accessor is constexpr-foldable, so the
|
||||
// algorithms compile to the same fully-unrolled/constant-bound code the old
|
||||
// hand-fixed MLP produced (verified against the RP2350 `.text` budget —
|
||||
// chokepoint B of docs/specs/plans/one-core-engine-refactor.md).
|
||||
//
|
||||
// Bit-parity contract: for identical shapes and seeds, MLPCore over
|
||||
// FixedStorage and DynamicStorage must produce bit-identical results — the
|
||||
// algorithm code is shared and the buffers are just memory. A ctest enforces
|
||||
// this (tests/cpp/test_mlp_storage_parity.cpp).
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <array>
|
||||
#include <cstddef>
|
||||
#include <span>
|
||||
|
||||
#include "../core/perf.hpp"
|
||||
|
||||
namespace nisps::ml {
|
||||
|
||||
inline constexpr std::size_t kMlpNumLayers = 4u;
|
||||
|
||||
template <std::size_t NIn,
|
||||
std::size_t NHidden1,
|
||||
std::size_t NHidden2,
|
||||
std::size_t NHidden3,
|
||||
std::size_t NOut,
|
||||
std::size_t NMaxExamples = 128u,
|
||||
std::size_t NMaxIterTrain = 4096u>
|
||||
class FixedStorage {
|
||||
public:
|
||||
static constexpr std::size_t kInput = NIn;
|
||||
static constexpr std::size_t kHidden1 = NHidden1;
|
||||
static constexpr std::size_t kHidden2 = NHidden2;
|
||||
static constexpr std::size_t kHidden3 = NHidden3;
|
||||
static constexpr std::size_t kOutput = NOut;
|
||||
static constexpr std::size_t kMaxExamples = NMaxExamples;
|
||||
static constexpr std::size_t kMaxIterTrain = NMaxIterTrain;
|
||||
static constexpr std::size_t kNumLayers = kMlpNumLayers;
|
||||
|
||||
static constexpr std::size_t weight_count() noexcept {
|
||||
return NIn * NHidden1 + NHidden1 * NHidden2 + NHidden2 * NHidden3 + NHidden3 * NOut
|
||||
+ NHidden1 + NHidden2 + NHidden3 + NOut;
|
||||
}
|
||||
|
||||
// ---- dims -----------------------------------------------------------
|
||||
static constexpr std::size_t n_in() noexcept { return NIn; }
|
||||
static constexpr std::size_t n_out() noexcept { return NOut; }
|
||||
static constexpr std::size_t max_examples() noexcept { return NMaxExamples; }
|
||||
static constexpr std::size_t max_iter_train() noexcept { return NMaxIterTrain; }
|
||||
|
||||
template <std::size_t L>
|
||||
static constexpr std::size_t fan_in_l() noexcept {
|
||||
static_assert(L < kNumLayers);
|
||||
if constexpr (L == 0u) return NIn;
|
||||
else if constexpr (L == 1u) return NHidden1;
|
||||
else if constexpr (L == 2u) return NHidden2;
|
||||
else return NHidden3;
|
||||
}
|
||||
template <std::size_t L>
|
||||
static constexpr std::size_t fan_out_l() noexcept {
|
||||
static_assert(L < kNumLayers);
|
||||
if constexpr (L == 0u) return NHidden1;
|
||||
else if constexpr (L == 1u) return NHidden2;
|
||||
else if constexpr (L == 2u) return NHidden3;
|
||||
else return NOut;
|
||||
}
|
||||
|
||||
// ---- per-layer buffers ------------------------------------------------
|
||||
template <std::size_t L> NISPS_FORCE_INLINE std::span<float> weights_l() noexcept {
|
||||
if constexpr (L == 0u) return w0_; else if constexpr (L == 1u) return w1_;
|
||||
else if constexpr (L == 2u) return w2_; else return w3_;
|
||||
}
|
||||
template <std::size_t L> NISPS_FORCE_INLINE std::span<const float> weights_l() const noexcept {
|
||||
if constexpr (L == 0u) return w0_; else if constexpr (L == 1u) return w1_;
|
||||
else if constexpr (L == 2u) return w2_; else return w3_;
|
||||
}
|
||||
template <std::size_t L> NISPS_FORCE_INLINE std::span<float> biases_l() noexcept {
|
||||
if constexpr (L == 0u) return b0_; else if constexpr (L == 1u) return b1_;
|
||||
else if constexpr (L == 2u) return b2_; else return b3_;
|
||||
}
|
||||
template <std::size_t L> NISPS_FORCE_INLINE std::span<const float> biases_l() const noexcept {
|
||||
if constexpr (L == 0u) return b0_; else if constexpr (L == 1u) return b1_;
|
||||
else if constexpr (L == 2u) return b2_; else return b3_;
|
||||
}
|
||||
template <std::size_t L> NISPS_FORCE_INLINE std::span<float> pre_act_l() noexcept {
|
||||
if constexpr (L == 0u) return pa0_; else if constexpr (L == 1u) return pa1_;
|
||||
else if constexpr (L == 2u) return pa2_; else return pa3_;
|
||||
}
|
||||
template <std::size_t L> NISPS_FORCE_INLINE std::span<float> act_l() noexcept {
|
||||
if constexpr (L == 0u) return a0_; else if constexpr (L == 1u) return a1_;
|
||||
else if constexpr (L == 2u) return a2_; else return a3_;
|
||||
}
|
||||
template <std::size_t L> NISPS_FORCE_INLINE std::span<const float> act_l() const noexcept {
|
||||
if constexpr (L == 0u) return a0_; else if constexpr (L == 1u) return a1_;
|
||||
else if constexpr (L == 2u) return a2_; else return a3_;
|
||||
}
|
||||
template <std::size_t L> NISPS_FORCE_INLINE std::span<float> grad_w_l() noexcept {
|
||||
if constexpr (L == 0u) return gw0_; else if constexpr (L == 1u) return gw1_;
|
||||
else if constexpr (L == 2u) return gw2_; else return gw3_;
|
||||
}
|
||||
template <std::size_t L> NISPS_FORCE_INLINE std::span<float> grad_b_l() noexcept {
|
||||
if constexpr (L == 0u) return gb0_; else if constexpr (L == 1u) return gb1_;
|
||||
else if constexpr (L == 2u) return gb2_; else return gb3_;
|
||||
}
|
||||
// Backprop scratch (delta into layer L's input), sized fan_in(L).
|
||||
template <std::size_t L> NISPS_FORCE_INLINE std::span<float> delta_l() noexcept {
|
||||
if constexpr (L == 0u) return d0_; else if constexpr (L == 1u) return d1_;
|
||||
else if constexpr (L == 2u) return d2_; else return d3_;
|
||||
}
|
||||
// Const-eval scratch (activation of layer L), sized fan_out(L). Mutable
|
||||
// so `eval_loss() const` can run the shared forward code without touching
|
||||
// the real activation caches.
|
||||
template <std::size_t L> NISPS_FORCE_INLINE std::span<float> eval_act_l() const noexcept {
|
||||
if constexpr (L == 0u) return e0_; else if constexpr (L == 1u) return e1_;
|
||||
else if constexpr (L == 2u) return e2_; else return e3_;
|
||||
}
|
||||
|
||||
// ---- global buffers ---------------------------------------------------
|
||||
NISPS_FORCE_INLINE std::span<float> input_buf() noexcept { return input_; }
|
||||
NISPS_FORCE_INLINE std::span<const float> input_buf() const noexcept { return input_; }
|
||||
NISPS_FORCE_INLINE std::span<float> output_buf() noexcept { return output_; }
|
||||
NISPS_FORCE_INLINE std::span<const float> output_buf() const noexcept { return output_; }
|
||||
NISPS_FORCE_INLINE std::span<float> ds_features() noexcept { return dsf_; }
|
||||
NISPS_FORCE_INLINE std::span<const float> ds_features() const noexcept { return dsf_; }
|
||||
NISPS_FORCE_INLINE std::span<float> ds_labels() noexcept { return dsl_; }
|
||||
NISPS_FORCE_INLINE std::span<const float> ds_labels() const noexcept { return dsl_; }
|
||||
NISPS_FORCE_INLINE std::span<float> flat_buf() noexcept { return flat_; }
|
||||
NISPS_FORCE_INLINE std::span<float> loss_hist_buf() noexcept { return lh_; }
|
||||
NISPS_FORCE_INLINE std::span<const float> loss_hist_buf() const noexcept { return lh_; }
|
||||
|
||||
private:
|
||||
std::array<float, NIn * NHidden1> w0_{};
|
||||
std::array<float, NHidden1 * NHidden2> w1_{};
|
||||
std::array<float, NHidden2 * NHidden3> w2_{};
|
||||
std::array<float, NHidden3 * NOut> w3_{};
|
||||
std::array<float, NHidden1> b0_{};
|
||||
std::array<float, NHidden2> b1_{};
|
||||
std::array<float, NHidden3> b2_{};
|
||||
std::array<float, NOut> b3_{};
|
||||
std::array<float, NHidden1> pa0_{};
|
||||
std::array<float, NHidden2> pa1_{};
|
||||
std::array<float, NHidden3> pa2_{};
|
||||
std::array<float, NOut> pa3_{};
|
||||
std::array<float, NHidden1> a0_{};
|
||||
std::array<float, NHidden2> a1_{};
|
||||
std::array<float, NHidden3> a2_{};
|
||||
std::array<float, NOut> a3_{};
|
||||
std::array<float, NIn * NHidden1> gw0_{};
|
||||
std::array<float, NHidden1 * NHidden2> gw1_{};
|
||||
std::array<float, NHidden2 * NHidden3> gw2_{};
|
||||
std::array<float, NHidden3 * NOut> gw3_{};
|
||||
std::array<float, NHidden1> gb0_{};
|
||||
std::array<float, NHidden2> gb1_{};
|
||||
std::array<float, NHidden3> gb2_{};
|
||||
std::array<float, NOut> gb3_{};
|
||||
std::array<float, NIn> d0_{};
|
||||
std::array<float, NHidden1> d1_{};
|
||||
std::array<float, NHidden2> d2_{};
|
||||
std::array<float, NHidden3> d3_{};
|
||||
mutable std::array<float, NHidden1> e0_{};
|
||||
mutable std::array<float, NHidden2> e1_{};
|
||||
mutable std::array<float, NHidden3> e2_{};
|
||||
mutable std::array<float, NOut> e3_{};
|
||||
|
||||
std::array<float, NIn> input_{};
|
||||
std::array<float, NOut> output_{};
|
||||
|
||||
std::array<float, NMaxExamples * NIn> dsf_{};
|
||||
std::array<float, NMaxExamples * NOut> dsl_{};
|
||||
|
||||
std::array<float, weight_count()> flat_{};
|
||||
std::array<float, NMaxIterTrain> lh_{};
|
||||
};
|
||||
|
||||
} // namespace nisps::ml
|
||||
|
|
@ -300,7 +300,7 @@ set_firmware_variant() {
|
|||
local selected
|
||||
|
||||
selected="$(choose_firmware_variant "$requested")"
|
||||
python - "$SKETCH_PATH" "$selected" <<'PY'
|
||||
"${PYTHON:-python3}" - "$SKETCH_PATH" "$selected" <<'PY'
|
||||
from pathlib import Path
|
||||
import re
|
||||
import sys
|
||||
|
|
|
|||
|
|
@ -16,6 +16,11 @@
|
|||
# - bare `new ` / `new(`
|
||||
# - malloc(
|
||||
# Files matching */tests/* are exempt — they are host-only.
|
||||
# SOLE allowlisted file: nisps/ml/dynamic_storage.hpp — the runtime-shaped
|
||||
# MLP storage (one arena allocation at construction). It is compile-time
|
||||
# excluded from RP2350 builds (#error under NISPS_TARGET_EMBEDDED); a
|
||||
# companion check below FAILS if that guard ever disappears, so heap can
|
||||
# not leak into firmware through the allowlist.
|
||||
#
|
||||
# 3. FAIL: `#include <Arduino.h>` anywhere under nisps/. The C++ core MUST
|
||||
# NOT pull in Arduino headers — those break the WASM build.
|
||||
|
|
@ -119,6 +124,7 @@ audit_heap_alloc() {
|
|||
hits=$(grep -REn "$pat" \
|
||||
--include='*.hpp' --include='*.cpp' \
|
||||
--exclude-dir=build --exclude-dir=tests \
|
||||
--exclude='dynamic_storage.hpp' \
|
||||
"${subdirs[@]}" 2>/dev/null \
|
||||
| grep -v ' *//' \
|
||||
|| true)
|
||||
|
|
@ -127,6 +133,14 @@ audit_heap_alloc() {
|
|||
echo "$hits" | sed 's/^/ /'
|
||||
fails=$((fails + 1))
|
||||
fi
|
||||
|
||||
# The allowlist above is only sound while dynamic_storage.hpp is
|
||||
# structurally excluded from embedded builds. Fail hard if the guard goes.
|
||||
local dyn="$NISPS_DIR/ml/dynamic_storage.hpp"
|
||||
if [[ -f "$dyn" ]] && ! grep -q 'NISPS_TARGET_EMBEDDED' "$dyn"; then
|
||||
echo "[lint-cpp] FAIL: $dyn lost its NISPS_TARGET_EMBEDDED #error guard"
|
||||
fails=$((fails + 1))
|
||||
fi
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
|
|
|
|||
161
tests/cpp/test_mlp_storage_parity.cpp
Normal file
161
tests/cpp/test_mlp_storage_parity.cpp
Normal file
|
|
@ -0,0 +1,161 @@
|
|||
// tests/cpp/test_mlp_storage_parity.cpp — FixedStorage vs DynamicStorage
|
||||
// bit-parity (one-core-engine-refactor P2 gate).
|
||||
//
|
||||
// For identical shapes and seeds, MLPCore over the two storage policies must
|
||||
// produce BIT-IDENTICAL results across the full surface: init, inference,
|
||||
// training, RL perturbation, diagnostics. Not 1e-5-near — memcmp-equal.
|
||||
|
||||
#include <cstdint>
|
||||
#include <cstring>
|
||||
#include <span>
|
||||
#include <vector>
|
||||
|
||||
#include "../../nisps/ml/dynamic_storage.hpp"
|
||||
#include "../../nisps/ml/mlp.hpp"
|
||||
#include "test_helpers.hpp"
|
||||
|
||||
namespace {
|
||||
|
||||
constexpr std::size_t kIn = 3u;
|
||||
constexpr std::size_t kH1 = 10u;
|
||||
constexpr std::size_t kH2 = 14u;
|
||||
constexpr std::size_t kH3 = 18u;
|
||||
constexpr std::size_t kOut = 7u;
|
||||
constexpr std::size_t kMaxEx = 16u;
|
||||
constexpr std::size_t kMaxIter = 64u;
|
||||
constexpr std::uint64_t kSeed = 0xC0FFEEu;
|
||||
|
||||
using FixedMLP = nisps::ml::MLP<kIn, kH1, kH2, kH3, kOut, kMaxEx, kMaxIter>;
|
||||
using DynamicMLP = nisps::ml::MLPCore<nisps::ml::DynamicStorage>;
|
||||
|
||||
DynamicMLP make_dynamic(std::uint64_t seed) {
|
||||
const std::size_t hidden[3] = {kH1, kH2, kH3};
|
||||
return DynamicMLP(seed, kIn, std::span<const std::size_t>(hidden), kOut, kMaxEx, kMaxIter);
|
||||
}
|
||||
|
||||
bool bit_equal(std::span<const float> a, std::span<const float> b) {
|
||||
if (a.size() != b.size()) return false;
|
||||
if (a.empty()) return true;
|
||||
return std::memcmp(a.data(), b.data(), a.size() * sizeof(float)) == 0;
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
// One scripted session driven through both storage models, checked
|
||||
// bit-exactly after every phase.
|
||||
NISPS_TEST(mlp_storage_parity_scripted_session) {
|
||||
FixedMLP fixed(kSeed);
|
||||
DynamicMLP dyn = make_dynamic(kSeed);
|
||||
|
||||
NISPS_ASSERT(dyn.valid());
|
||||
NISPS_ASSERT(fixed.weight_count() == dyn.weight_count());
|
||||
|
||||
// Construction (draw_weights(1.f) from the same seed).
|
||||
NISPS_EXPECT(bit_equal(fixed.get_weights(), dyn.get_weights()));
|
||||
|
||||
// Explicit draw at an interior spread.
|
||||
fixed.draw_weights(0.6f);
|
||||
dyn.draw_weights(0.6f);
|
||||
NISPS_EXPECT(bit_equal(fixed.get_weights(), dyn.get_weights()));
|
||||
|
||||
// Inference.
|
||||
const float probe_in[kIn] = {0.25f, 0.75f, 0.5f};
|
||||
for (std::size_t i = 0; i < kIn; ++i) {
|
||||
fixed.set_input(i, probe_in[i]);
|
||||
dyn.set_input(i, probe_in[i]);
|
||||
}
|
||||
fixed.process();
|
||||
dyn.process();
|
||||
NISPS_EXPECT(bit_equal(fixed.outputs(), dyn.outputs()));
|
||||
|
||||
// Dataset + training (enough examples to exercise the FIFO eviction).
|
||||
for (std::size_t e = 0; e < kMaxEx + 4u; ++e) {
|
||||
float feat[kIn];
|
||||
float lab[kOut];
|
||||
for (std::size_t i = 0; i < kIn; ++i) {
|
||||
feat[i] = 0.1f * static_cast<float>((e + i) % 10u);
|
||||
}
|
||||
for (std::size_t i = 0; i < kOut; ++i) {
|
||||
lab[i] = 0.05f * static_cast<float>((e * 3u + i) % 20u);
|
||||
}
|
||||
fixed.add_example(std::span<const float>(feat), std::span<const float>(lab));
|
||||
dyn.add_example(std::span<const float>(feat), std::span<const float>(lab));
|
||||
}
|
||||
NISPS_ASSERT(fixed.example_count() == dyn.example_count());
|
||||
|
||||
const float loss_f = fixed.train(0.5f, 40u, 0.0f);
|
||||
const float loss_d = dyn.train(0.5f, 40u, 0.0f);
|
||||
NISPS_EXPECT(std::memcmp(&loss_f, &loss_d, sizeof(float)) == 0);
|
||||
NISPS_EXPECT(bit_equal(fixed.get_weights(), dyn.get_weights()));
|
||||
NISPS_EXPECT(bit_equal(fixed.loss_history(), dyn.loss_history()));
|
||||
|
||||
// RL perturbation with a pin mask.
|
||||
std::uint8_t mask[kOut] = {};
|
||||
mask[2] = 1u;
|
||||
mask[5] = 1u;
|
||||
fixed.move_weights(0.3f, 0.4f, std::span<const std::uint8_t>(mask));
|
||||
dyn.move_weights(0.3f, 0.4f, std::span<const std::uint8_t>(mask));
|
||||
NISPS_EXPECT(bit_equal(fixed.get_weights(), dyn.get_weights()));
|
||||
|
||||
// Diagnostics.
|
||||
const float el_f = fixed.eval_loss();
|
||||
const float el_d = dyn.eval_loss();
|
||||
NISPS_EXPECT(std::memcmp(&el_f, &el_d, sizeof(float)) == 0);
|
||||
for (std::size_t l = 0; l < 4u; ++l) {
|
||||
const auto sf = fixed.layer_stats(l);
|
||||
const auto sd = dyn.layer_stats(l);
|
||||
NISPS_EXPECT(std::memcmp(&sf, &sd, sizeof(sf)) == 0);
|
||||
}
|
||||
|
||||
// set_weights round trip + infer_batch.
|
||||
{
|
||||
const auto wf = fixed.get_weights();
|
||||
std::vector<float> w(wf.begin(), wf.end());
|
||||
for (std::size_t i = 0; i < w.size(); i += 7u) w[i] += 0.125f;
|
||||
fixed.set_weights(w);
|
||||
dyn.set_weights(w);
|
||||
|
||||
const float pts[kIn * 3u] = {0.f, 0.f, 0.f,
|
||||
0.5f, 0.25f, 1.f,
|
||||
1.f, 1.f, 0.75f};
|
||||
float out_f[kOut * 3u];
|
||||
float out_d[kOut * 3u];
|
||||
fixed.infer_batch(std::span<const float>(pts), std::span<float>(out_f));
|
||||
dyn.infer_batch(std::span<const float>(pts), std::span<float>(out_d));
|
||||
NISPS_EXPECT(std::memcmp(out_f, out_d, sizeof(out_f)) == 0);
|
||||
}
|
||||
|
||||
// reset() re-draws from the (identically-advanced) RNG stream.
|
||||
fixed.reset();
|
||||
dyn.reset();
|
||||
NISPS_EXPECT(bit_equal(fixed.get_weights(), dyn.get_weights()));
|
||||
}
|
||||
|
||||
// Invalid dynamic construction stays inert (no crash, no UB).
|
||||
NISPS_TEST(mlp_dynamic_storage_invalid_dims_inert) {
|
||||
const std::size_t bad_hidden[2] = {4u, 4u};
|
||||
DynamicMLP bad(kSeed, kIn, std::span<const std::size_t>(bad_hidden), kOut);
|
||||
NISPS_ASSERT(!bad.valid());
|
||||
bad.process();
|
||||
bad.set_input(0u, 0.5f);
|
||||
NISPS_EXPECT(bad.train() == 0.f);
|
||||
NISPS_EXPECT(bad.get_weights().empty());
|
||||
NISPS_EXPECT(bad.eval_loss() == 0.f);
|
||||
}
|
||||
|
||||
// Moved-from dynamic instances stay inert; moved-to keeps working.
|
||||
NISPS_TEST(mlp_dynamic_storage_move_semantics) {
|
||||
DynamicMLP a = make_dynamic(kSeed);
|
||||
NISPS_ASSERT(a.valid());
|
||||
a.set_input(0u, 0.25f);
|
||||
a.process();
|
||||
|
||||
FixedMLP ref(kSeed);
|
||||
ref.set_input(0u, 0.25f);
|
||||
ref.process();
|
||||
|
||||
DynamicMLP b(static_cast<DynamicMLP&&>(a));
|
||||
NISPS_ASSERT(b.valid());
|
||||
b.process();
|
||||
NISPS_EXPECT(bit_equal(b.outputs(), ref.outputs()));
|
||||
}
|
||||
Loading…
Reference in a new issue