memlnaut-nisps/docs/specs/recon/findings-push-away-upstream-comparison.md
2026-07-25 16:14:35 +02:00

179 lines
10 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
kind: finding
date: 2026-07-25
immutable: true
---
# Findings — Manifold “Push away” vs upstream geometric dislike
_Read-only comparison, 2026-07-25. “Confirmed” claims cite source; “Inference” labels
interpretation. No implementation decision is made here._
## Scope and source identity
The hardware repository named in the question,
[`MusicallyEmbodiedML/MEMLNaut`](https://github.com/MusicallyEmbodiedML/MEMLNaut),
does not contain the active ML implementation. The firmware repository is
[`MusicallyEmbodiedML/MEMLNaut-NISPS`](https://github.com/MusicallyEmbodiedML/MEMLNaut-NISPS/tree/701f2d9f1b4698e0ddfa147193928489de12601f),
whose `src/memllib` submodule pins
[`MusicallyEmbodiedML/memllib` at `e291192`](https://github.com/MusicallyEmbodiedML/memllib/tree/e291192d8e4f2fca7b79670c4df9c2ec8bdf03cd).
The local read-only copies of upstream `InterfaceRL.hpp`, `.tpp`, and `.cpp` have Git
blob hashes `03e9255`, `9ec762b`, and `60204d1`, respectively—the same blobs returned
by GitHub for that pin. The comparison below is therefore against the exact upstream
source, not an approximation. Provenance is recorded locally in
`firmware/MEMLNaut-NISPS/lib/memllib/reference/README.md:1-15`.
## Confirmed behaviour
### The mental model is substantially right, but “different” needs a target
A negative verdict has no supervised label by itself. Upstream defines “different” as:
take the output that was heard, find the mean output of the four liked positions nearest
the current control input, and create a new target one unit farther from that liked
centroid in output space. With no likes—or a degenerate zero-length direction—it uses a
random direction. The target is clamped to `[0,1]`
([upstream `InterfaceRL.tpp:698-761`](https://github.com/MusicallyEmbodiedML/memllib/blob/e291192d8e4f2fca7b79670c4df9c2ec8bdf03cd/examples/InterfaceRL.tpp#L698-L761);
local mirror `firmware/MEMLNaut-NISPS/lib/memllib/reference/InterfaceRL.tpp:698-761`).
That target is then passed to `synthMapping.TrainBatch`. Inference subsequently reads
the same MLP through `synthMapping.GetOutput`. Therefore current upstream push-away is
**weight training on the mapping MLP**, not a separate rejection layer after inference
([upstream training call](https://github.com/MusicallyEmbodiedML/memllib/blob/e291192d8e4f2fca7b79670c4df9c2ec8bdf03cd/examples/InterfaceRL.tpp#L757-L761);
[upstream inference path](https://github.com/MusicallyEmbodiedML/memllib/blob/e291192d8e4f2fca7b79670c4df9c2ec8bdf03cd/examples/InterfaceRL.tpp#L869-L890)).
OU noise and `paramTransformHook` do exist after MLP inference, but are independent of
the dislike algorithm (`InterfaceRL.tpp:877-890`).
The Manifold path has the same basic semantics. The UI passes its current post-output-
pipeline vector to `FeedbackController.dislike`, which calls the WASM geometric-dislike
entry point (`manifold/src/console/ConsoleApp.tsx:509-535`;
`manifold/src/feedback/controller.ts:347-373`;
`manifold/src/engine/wasm-iml.ts:956-978`). The C++ controller computes the target and
calls `MLPCore::train_targets`, which runs forward propagation, backpropagation and one
RMSProp update on the same network weights (`nisps/ml/feedback.hpp:395-477`;
`nisps/ml/mlp.hpp:250-287`).
### The target maths now matches current upstream, but the training schedule does not
Current Manifold main now matches upstreams untapered target formula and constants:
`kGeometricPushScale=1.0`, `kNegLRBase=1.5`, no `/(1+distance)` taper
(`nisps/ml/geo_push.hpp:1-110`; upstream `InterfaceRL.hpp:406-414` and
`InterfaceRL.tpp:723-761`). It also now uses the upstream RMSProp update rather than
interpreting an upstream RMSProp learning rate as plain SGD
(`nisps/ml/training.hpp:1-92`).
The remaining major divergence is dose:
- Upstream stores the dislike at press time, then its main loop calls `optimise()` on
subsequent cycles (`InterfaceRL.tpp:38-54,186-232`).
- Every optimisation scans **all** live negatives and batch-trains them again
(`InterfaceRL.tpp:673-761`).
- Upstream computes one liked centroid around the **current live control input** and
applies it to every negative in that cycle, even if the user has moved away from the
original disliked position (`InterfaceRL.tpp:698-755`). Manifold instead computes the
centroid at the just-pressed negatives stored input (`nisps/ml/feedback.hpp:431-467`).
- A negative remains at full strength for 2500 ms, then expires; the number of updates
depends on the modes loop rate (`InterfaceRL.hpp:406-414`).
- Manifold collapses press and optimisation into **one synchronous
`train_targets` call for only the just-pressed negative**, then proportionally decays
stored negatives (`nisps/ml/feedback.hpp:395-477`; `nisps/ml/replay.hpp:165-185`).
There is no background/per-frame feedback optimiser.
Thus a Manifold click names a strongly displaced target, but takes only one step toward
it. Upstream keeps walking toward its target for the next 2.5 seconds. This is the most
direct explanation for a remaining perceptual strength difference.
As checked on 2026-07-25, the WASM served by
`https://meml.lnfinitemonkeys.org/next/nisps.wasm` has SHA-256
`d1c58a59517a00c6f51870ea1ec21194561b81058e22bbfb2e11de4af45c645a`, exactly matching
`manifold/public/nisps.wasm` on current main. The reported live behaviour therefore
cannot be explained by production still serving the pre-RMSProp or tapered binary.
### Upstream also cancels a nearby positive; Manifold does not
Upstreams default replay policy is `REPLACE_10_PERCENT`
(`InterfaceRL.hpp:404-405`). When a negative is stored, a positive within input-space
distance `0.10` is removed before the negative is added
(`InterfaceRL.tpp:904-929,950-979`). Manifolds replay method only deepens or adds a
negative and leaves positives intact (`nisps/ml/replay.hpp:102-121`). Manifold also
keeps liked examples in the separate MLP dataset
(`manifold/src/feedback/controller.ts:375-390`).
**Inference:** this is less about the first clicks amplitude than persistence. A later
positive training run can pull the mapping back toward a sound rejected near an
existing like, whereas upstream removes that local positive from its continuously
trained replay set.
### Current measured scale
On current main, the native behavioural benchmark at shape `2→16→16→16→8`, seed
`24301`, reports:
- one geometric dislike: at-point L2 movement `0.05335`;
- one legacy undirected Diffuse dislike: `0.22626`;
- repeated geometric dislikes: `0.05314` after 1, `0.33748` after 10, `1.09504`
after 100.
Commands:
```bash
scripts/bench-ml.sh --native-only --scenario A4_negative_once
scripts/bench-ml.sh --native-only --scenario D1_geo_anatomy
```
These numbers confirm that the current path is no longer inert, but also that a single
geometric press is still about 4.2× smaller than the legacy random-diffusion gesture in
this benchmark. They do not by themselves establish the right musical feel.
At Manifolds current default PAF shape (`4→10→10→14→33`), the same seeded scenarios
report one-click movement `0.09654` geometric versus `0.20788` Diffuse (about 2.2×
smaller), and geometric movement `0.55935` after ten presses. These are vector L2
distances across 33 parameters, so they establish that weights move; they do not prove
that the affected parameters produce a perceptually obvious timbral change.
The more revealing PAF-shape `A12_like_then_dislike` journey dislikes exactly where a
liked target was taught. With one update, distance from that rejected liked target
changes from `0.38291` to `0.36030` (`rejection_moved=-0.02261`): the mapping moves, but
slightly **toward** the particular target the user just rejected. Ten updates change the
distance to `0.67843` (`rejection_moved=+0.29552`). This is deterministic evidence that
one update is not sufficient to realise the user-facing semantic in an important
contradictory-feedback case; it also motivates testing upstreams nearby-like removal
separately.
## Why it likely felt weak
1. **Fixed today: optimiser mismatch.** Before `f57cddc`, the port used a tiny upstream
RMSProp learning rate inside plain SGD, reducing one press to roughly `5.3e-5`
movement.
2. **Fixed today: superseded push formula.** Before `ec31180`, the port halved the target
step, used one-third the negative-LR base, and tapered the step by distance—the exact
case upstream says made “no” ineffective.
3. **Still present: one update versus a time window of updates.** Manifold performs one
weight update per click; upstream replays all dislikes repeatedly for 2.5 seconds.
4. **Still present: highly asymmetric teaching dose.** A Manifold like trains the whole
positive dataset at `lr=1.0` for up to 1000 iterations, while a dislike takes one
roughly `0.0015` RMSProp step. Relative to the lurch caused by a like, correction
still feels small (`ALIGNMENT.md:66-84`).
5. **Potential persistence mismatch:** nearby likes are retained locally but removed by
upstreams default replay policy, so future positive retraining may partly undo the
rejection.
## Recommended next experiment
Do not start by increasing `geo_lr` blindly. That can make a click stronger but does not
answer whether the intended interaction is one discrete correction or a short-lived
repulsive constraint.
Add a deterministic “advance feedback by `dt`” core seam and benchmark three matched
variants from the same seeded prefix:
1. current one-shot update;
2. upstream-style replay of all negatives for 2500 ms, while separately testing whether
centroid lookup follows each stored negative or upstreams current live input;
3. a bounded discrete equivalent (for example 10/25/50 steps at press time) that avoids
wall-clock ownership in the core.
For each, report at-point movement, neighbourhood rings, global blast ratio, collateral
movement at liked positions, and like→dislike→like persistence. Then perform an
in-Manifold blind A/B at the real per-mode output arities and choose the smallest dose
that makes one rejection obvious without damaging liked regions. Test the nearby-like
removal policy as a separate axis rather than coupling it to dose.