memlnaut-nisps/docs/specs/recon/findings-push-away-upstream-comparison.md
2026-07-25 16:14:35 +02:00

10 KiB
Raw Permalink Blame History

kind date immutable
finding 2026-07-25 true

Findings — Manifold “Push away” vs upstream geometric dislike

Read-only comparison, 2026-07-25. “Confirmed” claims cite source; “Inference” labels interpretation. No implementation decision is made here.

Scope and source identity

The hardware repository named in the question, MusicallyEmbodiedML/MEMLNaut, does not contain the active ML implementation. The firmware repository is MusicallyEmbodiedML/MEMLNaut-NISPS, whose src/memllib submodule pins MusicallyEmbodiedML/memllib at e291192. The local read-only copies of upstream InterfaceRL.hpp, .tpp, and .cpp have Git blob hashes 03e9255, 9ec762b, and 60204d1, respectively—the same blobs returned by GitHub for that pin. The comparison below is therefore against the exact upstream source, not an approximation. Provenance is recorded locally in firmware/MEMLNaut-NISPS/lib/memllib/reference/README.md:1-15.

Confirmed behaviour

The mental model is substantially right, but “different” needs a target

A negative verdict has no supervised label by itself. Upstream defines “different” as: take the output that was heard, find the mean output of the four liked positions nearest the current control input, and create a new target one unit farther from that liked centroid in output space. With no likes—or a degenerate zero-length direction—it uses a random direction. The target is clamped to [0,1] (upstream InterfaceRL.tpp:698-761; local mirror firmware/MEMLNaut-NISPS/lib/memllib/reference/InterfaceRL.tpp:698-761).

That target is then passed to synthMapping.TrainBatch. Inference subsequently reads the same MLP through synthMapping.GetOutput. Therefore current upstream push-away is weight training on the mapping MLP, not a separate rejection layer after inference (upstream training call; upstream inference path). OU noise and paramTransformHook do exist after MLP inference, but are independent of the dislike algorithm (InterfaceRL.tpp:877-890).

The Manifold path has the same basic semantics. The UI passes its current post-output- pipeline vector to FeedbackController.dislike, which calls the WASM geometric-dislike entry point (manifold/src/console/ConsoleApp.tsx:509-535; manifold/src/feedback/controller.ts:347-373; manifold/src/engine/wasm-iml.ts:956-978). The C++ controller computes the target and calls MLPCore::train_targets, which runs forward propagation, backpropagation and one RMSProp update on the same network weights (nisps/ml/feedback.hpp:395-477; nisps/ml/mlp.hpp:250-287).

The target maths now matches current upstream, but the training schedule does not

Current Manifold main now matches upstreams untapered target formula and constants: kGeometricPushScale=1.0, kNegLRBase=1.5, no /(1+distance) taper (nisps/ml/geo_push.hpp:1-110; upstream InterfaceRL.hpp:406-414 and InterfaceRL.tpp:723-761). It also now uses the upstream RMSProp update rather than interpreting an upstream RMSProp learning rate as plain SGD (nisps/ml/training.hpp:1-92).

The remaining major divergence is dose:

  • Upstream stores the dislike at press time, then its main loop calls optimise() on subsequent cycles (InterfaceRL.tpp:38-54,186-232).
  • Every optimisation scans all live negatives and batch-trains them again (InterfaceRL.tpp:673-761).
  • Upstream computes one liked centroid around the current live control input and applies it to every negative in that cycle, even if the user has moved away from the original disliked position (InterfaceRL.tpp:698-755). Manifold instead computes the centroid at the just-pressed negatives stored input (nisps/ml/feedback.hpp:431-467).
  • A negative remains at full strength for 2500 ms, then expires; the number of updates depends on the modes loop rate (InterfaceRL.hpp:406-414).
  • Manifold collapses press and optimisation into one synchronous train_targets call for only the just-pressed negative, then proportionally decays stored negatives (nisps/ml/feedback.hpp:395-477; nisps/ml/replay.hpp:165-185). There is no background/per-frame feedback optimiser.

Thus a Manifold click names a strongly displaced target, but takes only one step toward it. Upstream keeps walking toward its target for the next 2.5 seconds. This is the most direct explanation for a remaining perceptual strength difference.

As checked on 2026-07-25, the WASM served by https://meml.lnfinitemonkeys.org/next/nisps.wasm has SHA-256 d1c58a59517a00c6f51870ea1ec21194561b81058e22bbfb2e11de4af45c645a, exactly matching manifold/public/nisps.wasm on current main. The reported live behaviour therefore cannot be explained by production still serving the pre-RMSProp or tapered binary.

Upstream also cancels a nearby positive; Manifold does not

Upstreams default replay policy is REPLACE_10_PERCENT (InterfaceRL.hpp:404-405). When a negative is stored, a positive within input-space distance 0.10 is removed before the negative is added (InterfaceRL.tpp:904-929,950-979). Manifolds replay method only deepens or adds a negative and leaves positives intact (nisps/ml/replay.hpp:102-121). Manifold also keeps liked examples in the separate MLP dataset (manifold/src/feedback/controller.ts:375-390).

Inference: this is less about the first clicks amplitude than persistence. A later positive training run can pull the mapping back toward a sound rejected near an existing like, whereas upstream removes that local positive from its continuously trained replay set.

Current measured scale

On current main, the native behavioural benchmark at shape 2→16→16→16→8, seed 24301, reports:

  • one geometric dislike: at-point L2 movement 0.05335;
  • one legacy undirected Diffuse dislike: 0.22626;
  • repeated geometric dislikes: 0.05314 after 1, 0.33748 after 10, 1.09504 after 100.

Commands:

scripts/bench-ml.sh --native-only --scenario A4_negative_once
scripts/bench-ml.sh --native-only --scenario D1_geo_anatomy

These numbers confirm that the current path is no longer inert, but also that a single geometric press is still about 4.2× smaller than the legacy random-diffusion gesture in this benchmark. They do not by themselves establish the right musical feel.

At Manifolds current default PAF shape (4→10→10→14→33), the same seeded scenarios report one-click movement 0.09654 geometric versus 0.20788 Diffuse (about 2.2× smaller), and geometric movement 0.55935 after ten presses. These are vector L2 distances across 33 parameters, so they establish that weights move; they do not prove that the affected parameters produce a perceptually obvious timbral change.

The more revealing PAF-shape A12_like_then_dislike journey dislikes exactly where a liked target was taught. With one update, distance from that rejected liked target changes from 0.38291 to 0.36030 (rejection_moved=-0.02261): the mapping moves, but slightly toward the particular target the user just rejected. Ten updates change the distance to 0.67843 (rejection_moved=+0.29552). This is deterministic evidence that one update is not sufficient to realise the user-facing semantic in an important contradictory-feedback case; it also motivates testing upstreams nearby-like removal separately.

Why it likely felt weak

  1. Fixed today: optimiser mismatch. Before f57cddc, the port used a tiny upstream RMSProp learning rate inside plain SGD, reducing one press to roughly 5.3e-5 movement.
  2. Fixed today: superseded push formula. Before ec31180, the port halved the target step, used one-third the negative-LR base, and tapered the step by distance—the exact case upstream says made “no” ineffective.
  3. Still present: one update versus a time window of updates. Manifold performs one weight update per click; upstream replays all dislikes repeatedly for 2.5 seconds.
  4. Still present: highly asymmetric teaching dose. A Manifold like trains the whole positive dataset at lr=1.0 for up to 1000 iterations, while a dislike takes one roughly 0.0015 RMSProp step. Relative to the lurch caused by a like, correction still feels small (ALIGNMENT.md:66-84).
  5. Potential persistence mismatch: nearby likes are retained locally but removed by upstreams default replay policy, so future positive retraining may partly undo the rejection.

Do not start by increasing geo_lr blindly. That can make a click stronger but does not answer whether the intended interaction is one discrete correction or a short-lived repulsive constraint.

Add a deterministic “advance feedback by dt” core seam and benchmark three matched variants from the same seeded prefix:

  1. current one-shot update;
  2. upstream-style replay of all negatives for 2500 ms, while separately testing whether centroid lookup follows each stored negative or upstreams current live input;
  3. a bounded discrete equivalent (for example 10/25/50 steps at press time) that avoids wall-clock ownership in the core.

For each, report at-point movement, neighbourhood rings, global blast ratio, collateral movement at liked positions, and like→dislike→like persistence. Then perform an in-Manifold blind A/B at the real per-mode output arities and choose the smallest dose that makes one rejection obvious without damaging liked regions. Test the nearby-like removal policy as a separate axis rather than coupling it to dose.