Technical note + data release accompanying Valid and False Snapping in EML Expression Trees: The Basin Selection Problem (DOI: 10.5281/zenodo.19790799; v2.3 at time of writing). The companion paper showed that three-phase temperature annealing (Adam on MAE, entropy-penalty ramp, temperature anneal) solves selector commitment in balanced EML expression trees (every selector reaches a simplex vertex) but does not solve basin selection: the snapped expression is often the wrong symbolic form. Under the paper's strict validity criterion (vertex commitment plus post-snap MAE < 0.01), blind initialization recovers the correct form in only 25–35% of runs on the difficult cells: exp(x) at low depth, where the competing basin eml(x,x) (post-snap MAE 0.688) captures most seeds, and ln(x) at over-representational depth 5, where structurally near-miss forms dominate. Contribution This note shows that two cheap, symbolic-aware interventions at initialization rescue these cells completely, with no change to the training schedule, loss, architecture, or validity criterion. 1. Targeted Warm-Start (initialize_to_target) Selector logits are biased toward the known-good form, plus Gaussian exploration noise (0.4). Cell Result exp d=3 20/20 valid recovery, every run ending at eml(x,1) ln d=5 20/20 valid recovery, every run ending at eml(1,eml(eml(1,x),1)) Post-snap loss ~1e-8 in all cases. 2. Top-Aligned Curriculum (grow_from_shallow + reanneal_extra_capacity) Train at the representational depth, embed the trained solution top-aligned into the deeper tree so unused capacity dangles below as constant-biased branches, briefly re-anneal only the extra selectors (200 epochs, embedded block frozen), then fine-tune. Top alignment is the load-bearing choice: a balanced EML gate cannot act as an identity forwarder (no a, b ∈ {1, x, f} gives exp(a) − ln(b) = f), so the naive bottom-aligned extension would embed exp(ln(x)) = x before training. The curriculum assumes only that the target is learnable at its representational depth — not that its deep-tree form is known — and also reaches 20/20 on both cells. A replay audit confirms every ln curriculum run used the genuine grow_from_shallow path rather than the driver's fallback. Blind Baselines (identical deterministic seeds) Cell Valid Recovery exp d=3 7/20 ln d=5 5/20 The results corroborate the companion paper's reframing of symbolic recovery as basin selection rather than commitment or representational power, and align with Odrzywolek's independent warm-start evidence (arXiv:2603.21852, SI Table S7) that correct solutions are stable attractors. Files in This Record File Description Note (PDF) Full technical note Figure 1 Grouped recovery-rate bars, both cells × three init modes basin_warmstart_v2.4_postfix.csv 120 rows: blind + warm controls basin_warmstart_v2.5_unbalanced.csv 120 rows: adds the top-aligned curriculum; blind/warm rows match v2.4_postfix run-for-run under identical seeding Both CSVs carry deterministic seeds and the auditable pretrain_form column. Code (eml_layer_v2.py, basin_warmstart.py) and exact reproduction commands are in the eml-ice40-cybenko repository, archived at 10.5281/zenodo.19736075. The note includes an AI Utilization Statement.
symbolic regression · EML operator · neural symbolic · expression trees · basin selection · warm-start initialization · curriculum learning · temperature annealing · elementary functions